Compare / head-to-head
Grok 4.6vs
Muse Spark 1.2
Grok 4.6 leads 11 of 20 shared benchmarks. Muse Spark 1.2 is 1.5x cheaper per token. Muse Spark 1.2 has the larger context window (1.0M tokens).
Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.
Shared benchmarks
11 – 8
Grok 4.6 leads · 1 tied
Cheaper per token
Muse Spark 1.2
1.5x cheaper, input + output
Larger context
Muse Spark 1.2
1.0M tokens
Providers
2 providers
xAI · Meta
Grok 4.6
xAI · released 2026-08-12
textvision
Muse Spark 1.2
Meta · released 2026-08-05
textvisionvideoaudio
Quality
Benchmark matrix
20 shared · 18 only Grok 4.6 · 3 only Muse Spark 1.2| Benchmark | Grok 4.6 | Muse Spark 1.2 | Δ | Edge |
|---|---|---|---|---|
| Reported by both · 20 | ||||
| LMArena Elo | 1453.8 ↗ | 1493.5 ↗ | -39.7 | Muse Spark 1.2 |
| APEX-Agents (Mercor) | 65.3% ↗ | 36.4% ↗ | +28.9 pt | Grok 4.6 |
| DeepSWE v1.1 | 65.9% ↗ | 59.3% ↗ | +6.6 pt | Grok 4.6 |
| DeepSWE v1.1 (Datacurve) | 67.5% ↗ | 54.9% ↗ | +12.6 pt | Grok 4.6 |
| FrontierSWE V2 (Proximal Labs) | 25.3% ↗ | 12% ↗ | +13.3 pt | Grok 4.6 |
| LiveBench Agentic Coding (LiveBench) | 57% ↗ | 57.6% ↗ | -0.6 pt | Muse Spark 1.2 |
| LiveBench Coding (LiveBench) | 76.8% ↗ | 77.5% ↗ | -0.7 pt | Muse Spark 1.2 |
| LiveBench Data Analysis (LiveBench) | 73.9% ↗ | 76.5% ↗ | -2.6 pt | Muse Spark 1.2 |
| LiveBench Instruction Following (LiveBench) | 71.9% ↗ | 74.3% ↗ | -2.4 pt | Muse Spark 1.2 |
| LiveBench Language (LiveBench) | 83.7% ↗ | 78.6% ↗ | +5.1 pt | Grok 4.6 |
| LiveBench Mathematics (LiveBench) | 92.6% ↗ | 91.2% ↗ | +1.4 pt | Grok 4.6 |
| LiveBench Reasoning (LiveBench) | 90.5% ↗ | 90% ↗ | +0.5 pt | Grok 4.6 |
| LMArena Agent (LMArena) | 0.0128 ↗ | -0.0327 ↗ | 0 | Tie |
| LMArena Vision (LMArena) | 1263.5 ↗ | 1292.8 ↗ | -29.3 | Muse Spark 1.2 |
| LMArena WebDev (LMArena) | 1619.5 ↗ | 1531.8 ↗ | +87.7 | Grok 4.6 |
| SAGE (Vals AI) | 28.9% ↗ | 47.7% ↗ | -18.8 pt | Muse Spark 1.2 |
| SimpleBench (SimpleBench) | 75.9% ↗ | 74.5% ↗ | +1.4 pt | Grok 4.6 |
| SimpleQA Verified | 49.3% ↗epoch run | 60.3% ↗epoch run | -11 pt | Muse Spark 1.2 |
| Terminal-Bench 4.0 (Vals AI) | 17.2% ↗ | 6.1% ↗ | +11.1 pt | Grok 4.6 |
| WeirdML (Håvard Tveit Ihle) | 67.3% ↗ | 60.3% ↗ | +7 pt | Grok 4.6 |
| Only Grok 4.6 reports · 18 | ||||
| GPQA Diamond | 94% ↗epoch run | not reported | — | — |
| AA-Briefcase | 1577 ↗ | not reported | — | — |
| APEX-Agents | 57.5% ↗ | not reported | — | — |
| APEX-SWE | 56.4% ↗ | not reported | — | — |
| Artificial Analysis Intelligence Index | 61 ↗ | not reported | — | — |
| Chess Puzzles (Epoch AI run) | 40% ↗epoch run | not reported | — | — |
| CursorBench 3.2 | 69.9% ↗ | not reported | — | — |
| EBR-bench (Epoch AI run) | 30.5% ↗epoch run | not reported | — | — |
| FrontierCode v1.1 (Extended) | 61.3% ↗ | not reported | — | — |
| FrontierMath Tier 4 v2 (Epoch AI run) | 31.7% ↗epoch run | not reported | — | — |
| FrontierMath Tiers 1-3 v2 (Epoch AI run) | 66% ↗epoch run | not reported | — | — |
| Furniture Assembly (Epoch AI run) | 40% ↗epoch run | not reported | — | — |
| GDPVal-AA v2 | 1753 ↗ | not reported | — | — |
| Harvey LAB (Vals) | 15.8% ↗ | not reported | — | — |
| Mystery Game Puzzles (Epoch AI run) | 34% ↗epoch run | not reported | — | — |
| OTIS Mock AIME 2024-2025 (Epoch AI run) | 99.2% ↗epoch run | not reported | — | — |
| Terminal-Bench 3.0 | 26% ↗ | not reported | — | — |
| Vending-Bench 2 (Andon Labs) | 9047.03 ↗ | not reported | — | — |
| Only Muse Spark 1.2 reports · 3 | ||||
| MCP Atlas | not reported | 90.3% ↗ | — | — |
| Terminal-Bench 2.1 | not reported | 82.9% ↗ | — | — |
| Toolathlon-Verified (HKUST) | not reported | 75.9% ↗ | — | — |
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Grok 4.6 minus Muse Spark 1.2 in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing
Side by side
Official pricing pages and model cards| Spec | Grok 4.6 | Muse Spark 1.2 | Edge |
|---|---|---|---|
| Context window | 500K tokens | 1.0M tokens | Muse Spark 1.2 |
| Max output | — | — | — |
| Input price / 1M | $2 ↗ | $1.25 ↗ | Muse Spark 1.2 |
| Output price / 1M | $6 ↗ | $4.25 ↗ | Muse Spark 1.2 |
| Cached input / 1M | $0.5 ↗ | $0.15 ↗ | Muse Spark 1.2 |
| Input + output / 1M Lower is cheaper. List prices; batch, tool and regional fees excluded. | $8 | $5.5 | Muse Spark 1.2 |
| Modalities | text · vision | text · vision · video · audio | Muse Spark 1.2 |
| Released | 2026-08-12 | 2026-08-05 | — |
| Cited benchmark scores | 43 | 23 | — |
Reliability
Provider status
Live from /statusMore matchups
Grok 4.6 vs …
Models sharing the most benchmarksMore matchups
Muse Spark 1.2 vs …
Models sharing the most benchmarksBuilt by Respan
Which one wins on your data?
Public benchmarks are a starting point. Run Grok 4.6 and Muse Spark 1.2 on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.