Compare / head-to-head

Claude Sonnet 5.5vsMuse Spark 1.2

Claude Sonnet 5.5 leads 12 of 17 shared benchmarks. Muse Spark 1.2 is 2.2x cheaper per token. Muse Spark 1.2 has the larger context window (1.0M tokens).

Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.

Shared benchmarks
12 – 5
Claude Sonnet 5.5 leads
Cheaper per token
Muse Spark 1.2
2.2x cheaper, input + output
Larger context
Muse Spark 1.2
1.0M tokens
Providers
2 providers
Anthropic · Meta
Claude Sonnet 5.5
Anthropic · released 2026-09-28
textvision
Context
1M
Max out
128K
Input /1M
$2 ↗
Output /1M
$10 ↗
Cached /1M
$0.2
Scores
29 · 10 core
Muse Spark 1.2
Meta · released 2026-08-05
textvisionvideoaudio
Context
1.0M
Max out
—
Input /1M
$1.25 ↗
Output /1M
$4.25 ↗
Cached /1M
$0.15
Scores
23 · 5 core
Quality

Benchmark matrix

BenchmarkClaude Sonnet 5.5Muse Spark 1.2ΔEdge
Reported by both · 17
LMArena Elo1471 ↗1493.5 ↗-22.5Muse Spark 1.2
APEX-Agents (Mercor)75.5% ↗36.4% ↗+39.1 ptClaude Sonnet 5.5
FrontierSWE V2 (Proximal Labs)61.9% ↗12% ↗+49.9 ptClaude Sonnet 5.5
LiveBench Agentic Coding (LiveBench)56.3% ↗57.6% ↗-1.3 ptMuse Spark 1.2
LiveBench Coding (LiveBench)91.4% ↗77.5% ↗+13.9 ptClaude Sonnet 5.5
LiveBench Data Analysis (LiveBench)78.6% ↗76.5% ↗+2.1 ptClaude Sonnet 5.5
LiveBench Instruction Following (LiveBench)70.5% ↗74.3% ↗-3.8 ptMuse Spark 1.2
LiveBench Language (LiveBench)83.4% ↗78.6% ↗+4.8 ptClaude Sonnet 5.5
LiveBench Mathematics (LiveBench)96.7% ↗91.2% ↗+5.5 ptClaude Sonnet 5.5
LiveBench Reasoning (LiveBench)91.6% ↗90% ↗+1.6 ptClaude Sonnet 5.5
LMArena Agent (LMArena)0.1252 ↗-0.0327 ↗+0.2Claude Sonnet 5.5
LMArena Vision (LMArena)1268.3 ↗1292.8 ↗-24.5Muse Spark 1.2
LMArena WebDev (LMArena)1786.3 ↗1531.8 ↗+254.5Claude Sonnet 5.5
SAGE (Vals AI)51.8% ↗47.7% ↗+4.1 ptClaude Sonnet 5.5
SimpleBench (SimpleBench)75.9% ↗74.5% ↗+1.4 ptClaude Sonnet 5.5
SimpleQA Verified46.5% ↗epoch run60.3% ↗epoch run-13.8 ptMuse Spark 1.2
Terminal-Bench 4.0 (Vals AI)64.1% ↗6.1% ↗+58 ptClaude Sonnet 5.5
Only Claude Sonnet 5.5 reports · 12
GPQA Diamond95.6% ↗epoch runnot reported——
Chartography (no tools)61.6% ↗not reported——
CursorBench 4.055.5% ↗not reported——
FrontierCode v1.1 (Main)46.2% ↗not reported——
FrontierMath Tier 4 v2 (Epoch AI run)80.5% ↗epoch runnot reported——
FrontierMath Tiers 1-3 v2 (Epoch AI run)88.8% ↗epoch runnot reported——
Furniture Assembly (Epoch AI run)75% ↗epoch runnot reported——
Humanity's Last Exam (with tools)64.5% ↗not reported——
Mystery Game Puzzles (Epoch AI run)65% ↗epoch runnot reported——
OSWorld 2.1 (partial)80.1% ↗not reported——
OTIS Mock AIME 2024-2025 (Epoch AI run)100% ↗epoch runnot reported——
Terminal-Bench 4.070.6% ↗not reported——
Only Muse Spark 1.2 reports · 6
DeepSWE v1.1not reported59.3% ↗——
DeepSWE v1.1 (Datacurve)not reported54.9% ↗——
MCP Atlasnot reported90.3% ↗——
Terminal-Bench 2.1not reported82.9% ↗——
Toolathlon-Verified (HKUST)not reported75.9% ↗——
WeirdML (Håvard Tveit Ihle)not reported60.3% ↗——
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Claude Sonnet 5.5 minus Muse Spark 1.2 in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing

Side by side

SpecClaude Sonnet 5.5Muse Spark 1.2Edge
Context window1M tokens1.0M tokensMuse Spark 1.2
Max output128K tokens——
Input price / 1M$2 ↗$1.25 ↗Muse Spark 1.2
Output price / 1M$10 ↗$4.25 ↗Muse Spark 1.2
Cached input / 1M$0.2 ↗$0.15 ↗Muse Spark 1.2
Input + output / 1M
Lower is cheaper. List prices; batch, tool and regional fees excluded.
$12$5.5Muse Spark 1.2
Modalitiestext · visiontext · vision · video · audioMuse Spark 1.2
Released2026-09-282026-08-05—
Cited benchmark scores2923—
Reliability

Provider status

All providers →
More matchups

Claude Sonnet 5.5 vs …

More matchups

Muse Spark 1.2 vs …

Built by Respan
Which one wins on your data?

Public benchmarks are a starting point. Run Claude Sonnet 5.5 and Muse Spark 1.2 on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.