Compare / head-to-head

Grok 4.6vsMuse Spark 1.2

Grok 4.6 leads 11 of 20 shared benchmarks. Muse Spark 1.2 is 1.5x cheaper per token. Muse Spark 1.2 has the larger context window (1.0M tokens).

Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.

Shared benchmarks
11 – 8
Grok 4.6 leads · 1 tied
Cheaper per token
Muse Spark 1.2
1.5x cheaper, input + output
Larger context
Muse Spark 1.2
1.0M tokens
Providers
2 providers
xAI · Meta
Grok 4.6
xAI · released 2026-08-12
textvision
Context
500K
Max out
—
Input /1M
$2 ↗
Output /1M
$6 ↗
Cached /1M
$0.5
Scores
43 · 9 core
Muse Spark 1.2
Meta · released 2026-08-05
textvisionvideoaudio
Context
1.0M
Max out
—
Input /1M
$1.25 ↗
Output /1M
$4.25 ↗
Cached /1M
$0.15
Scores
23 · 5 core
Quality

Benchmark matrix

BenchmarkGrok 4.6Muse Spark 1.2ΔEdge
Reported by both · 20
LMArena Elo1453.8 ↗1493.5 ↗-39.7Muse Spark 1.2
APEX-Agents (Mercor)65.3% ↗36.4% ↗+28.9 ptGrok 4.6
DeepSWE v1.165.9% ↗59.3% ↗+6.6 ptGrok 4.6
DeepSWE v1.1 (Datacurve)67.5% ↗54.9% ↗+12.6 ptGrok 4.6
FrontierSWE V2 (Proximal Labs)25.3% ↗12% ↗+13.3 ptGrok 4.6
LiveBench Agentic Coding (LiveBench)57% ↗57.6% ↗-0.6 ptMuse Spark 1.2
LiveBench Coding (LiveBench)76.8% ↗77.5% ↗-0.7 ptMuse Spark 1.2
LiveBench Data Analysis (LiveBench)73.9% ↗76.5% ↗-2.6 ptMuse Spark 1.2
LiveBench Instruction Following (LiveBench)71.9% ↗74.3% ↗-2.4 ptMuse Spark 1.2
LiveBench Language (LiveBench)83.7% ↗78.6% ↗+5.1 ptGrok 4.6
LiveBench Mathematics (LiveBench)92.6% ↗91.2% ↗+1.4 ptGrok 4.6
LiveBench Reasoning (LiveBench)90.5% ↗90% ↗+0.5 ptGrok 4.6
LMArena Agent (LMArena)0.0128 ↗-0.0327 ↗0Tie
LMArena Vision (LMArena)1263.5 ↗1292.8 ↗-29.3Muse Spark 1.2
LMArena WebDev (LMArena)1619.5 ↗1531.8 ↗+87.7Grok 4.6
SAGE (Vals AI)28.9% ↗47.7% ↗-18.8 ptMuse Spark 1.2
SimpleBench (SimpleBench)75.9% ↗74.5% ↗+1.4 ptGrok 4.6
SimpleQA Verified49.3% ↗epoch run60.3% ↗epoch run-11 ptMuse Spark 1.2
Terminal-Bench 4.0 (Vals AI)17.2% ↗6.1% ↗+11.1 ptGrok 4.6
WeirdML (Håvard Tveit Ihle)67.3% ↗60.3% ↗+7 ptGrok 4.6
Only Grok 4.6 reports · 18
GPQA Diamond94% ↗epoch runnot reported——
AA-Briefcase1577 ↗not reported——
APEX-Agents57.5% ↗not reported——
APEX-SWE56.4% ↗not reported——
Artificial Analysis Intelligence Index61 ↗not reported——
Chess Puzzles (Epoch AI run)40% ↗epoch runnot reported——
CursorBench 3.269.9% ↗not reported——
EBR-bench (Epoch AI run)30.5% ↗epoch runnot reported——
FrontierCode v1.1 (Extended)61.3% ↗not reported——
FrontierMath Tier 4 v2 (Epoch AI run)31.7% ↗epoch runnot reported——
FrontierMath Tiers 1-3 v2 (Epoch AI run)66% ↗epoch runnot reported——
Furniture Assembly (Epoch AI run)40% ↗epoch runnot reported——
GDPVal-AA v21753 ↗not reported——
Harvey LAB (Vals)15.8% ↗not reported——
Mystery Game Puzzles (Epoch AI run)34% ↗epoch runnot reported——
OTIS Mock AIME 2024-2025 (Epoch AI run)99.2% ↗epoch runnot reported——
Terminal-Bench 3.026% ↗not reported——
Vending-Bench 2 (Andon Labs)9047.03 ↗not reported——
Only Muse Spark 1.2 reports · 3
MCP Atlasnot reported90.3% ↗——
Terminal-Bench 2.1not reported82.9% ↗——
Toolathlon-Verified (HKUST)not reported75.9% ↗——
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Grok 4.6 minus Muse Spark 1.2 in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing

Side by side

SpecGrok 4.6Muse Spark 1.2Edge
Context window500K tokens1.0M tokensMuse Spark 1.2
Max output———
Input price / 1M$2 ↗$1.25 ↗Muse Spark 1.2
Output price / 1M$6 ↗$4.25 ↗Muse Spark 1.2
Cached input / 1M$0.5 ↗$0.15 ↗Muse Spark 1.2
Input + output / 1M
Lower is cheaper. List prices; batch, tool and regional fees excluded.
$8$5.5Muse Spark 1.2
Modalitiestext · visiontext · vision · video · audioMuse Spark 1.2
Released2026-08-122026-08-05—
Cited benchmark scores4323—
Reliability

Provider status

All providers →
More matchups

Grok 4.6 vs …

More matchups

Muse Spark 1.2 vs …

Built by Respan
Which one wins on your data?

Public benchmarks are a starting point. Run Grok 4.6 and Muse Spark 1.2 on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.