Compare / head-to-head

Claude Opus 4.5vsMuse Spark 1.3

Muse Spark 1.3 leads 15 of 15 shared benchmarks. Muse Spark 1.3 is 5.5x cheaper per token. Muse Spark 1.3 has the larger context window (1.0M tokens).

Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.

Shared benchmarks
0 – 15
Muse Spark 1.3 leads
Cheaper per token
Muse Spark 1.3
5.5x cheaper, input + output
Larger context
Muse Spark 1.3
1.0M tokens
Providers
2 providers
Anthropic · Meta
Claude Opus 4.5
Anthropic · released 2025-11-24
textvision
Context
200K
Max out
64K
Input /1M
$5 ↗
Output /1M
$25 ↗
Cached /1M
$0.5
Scores
39 · 9 core
Muse Spark 1.3
Meta · released 2026-09-02
textvisionvideo
Context
1.0M
Max out
—
Input /1M
$1.25 ↗
Output /1M
$4.25 ↗
Cached /1M
$0.15
Scores
29 · 7 core
Quality

Benchmark matrix

BenchmarkClaude Opus 4.5Muse Spark 1.3ΔEdge
Reported by both · 15
LMArena Elo1473.6 ↗1494.3 ↗-20.7Muse Spark 1.3
Chess Puzzles (Epoch AI run)12% ↗epoch run38% ↗epoch run-26 ptMuse Spark 1.3
FrontierMath Tier 4 v2 (Epoch AI run)4.9% ↗epoch run46.3% ↗epoch run-41.4 ptMuse Spark 1.3
FrontierMath Tiers 1-3 v2 (Epoch AI run)34.4% ↗epoch run74.4% ↗epoch run-40 ptMuse Spark 1.3
LiveBench Agentic Coding (LiveBench)39.7% ↗64.1% ↗-24.4 ptMuse Spark 1.3
LiveBench Coding (LiveBench)79.7% ↗81.1% ↗-1.4 ptMuse Spark 1.3
LiveBench Data Analysis (LiveBench)74.4% ↗79.6% ↗-5.2 ptMuse Spark 1.3
LiveBench Instruction Following (LiveBench)62.5% ↗78% ↗-15.5 ptMuse Spark 1.3
LiveBench Language (LiveBench)81.3% ↗82.8% ↗-1.5 ptMuse Spark 1.3
LiveBench Mathematics (LiveBench)90.4% ↗95.9% ↗-5.5 ptMuse Spark 1.3
LiveBench Reasoning (LiveBench)80.1% ↗89.7% ↗-9.6 ptMuse Spark 1.3
LMArena WebDev (LMArena)1493.3 ↗1656.6 ↗-163.3Muse Spark 1.3
Mystery Game Puzzles (Epoch AI run)22% ↗epoch run25% ↗epoch run-3 ptMuse Spark 1.3
OTIS Mock AIME 2024-2025 (Epoch AI run)86.1% ↗epoch run99.2% ↗epoch run-13.1 ptMuse Spark 1.3
SimpleBench (SimpleBench)62% ↗81.8% ↗-19.8 ptMuse Spark 1.3
Only Claude Opus 4.5 reports · 23
SWE-bench Verified80.9% ↗not reported——
GPQA Diamond87% ↗not reported——
ARC-AGI-2 (Verified)37.6% ↗not reported——
BALROG (BALROG)43.5% ↗not reported——
EBR-bench (Epoch AI run)14.3% ↗epoch runnot reported——
Furniture Assembly (Epoch AI run)28.3% ↗epoch runnot reported——
GSO Opt@1 (GSO)24.5% ↗not reported——
MCP Atlas62.3% ↗not reported——
MMMLU90.8% ↗not reported——
MMMU (validation)80.7% ↗not reported——
OSWorld66.3% ↗not reported——
SAGE (Vals AI)52.1% ↗not reported——
SimpleQA Verified45.7% ↗epoch runnot reported——
SWE-bench Verified (Epoch AI run)76.7% ↗epoch runnot reported——
tau2-bench Airline (Sierra)84% ↗not reported——
tau2-bench Banking Knowledge (Sierra)24.7% ↗not reported——
tau2-bench Retail (Sierra)79.6% ↗not reported——
tau2-bench Telecom (Sierra)92.3% ↗not reported——
Terminal-Bench 2.059.3% ↗not reported——
Vending-Bench 2 (Andon Labs)4967.06 ↗not reported——
WeirdML (Håvard Tveit Ihle)63.7% ↗not reported——
τ2-bench (Retail)88.9% ↗not reported——
τ2-bench (Telecom)98.2% ↗not reported——
Only Muse Spark 1.3 reports · 14
APEX-Agents (Mercor)not reported57.8% ↗——
AutomationBenchnot reported49.6% ↗——
DeepSearchQAnot reported90.3% ↗——
DeepSWE v1.1not reported75.4% ↗——
GDPval-AA v2not reported1754 ↗——
JobBenchnot reported64.9% ↗——
LMArena Agent (LMArena)not reported0.04 ↗——
LMArena Vision (LMArena)not reported1290.2 ↗——
MRCR 256K-512Knot reported98.5% ↗——
MRCR 512K-1Mnot reported98.1% ↗——
OSWorld 2.0 (partial)not reported66.9% ↗——
SWE-Atlas Codebase QnAnot reported59.4% ↗——
Terminal-Bench 2.1not reported88.8% ↗——
Terminal-Bench 4.0 (Vals AI)not reported24.7% ↗——
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Claude Opus 4.5 minus Muse Spark 1.3 in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing

Side by side

SpecClaude Opus 4.5Muse Spark 1.3Edge
Context window200K tokens1.0M tokensMuse Spark 1.3
Max output64K tokens——
Input price / 1M$5 ↗$1.25 ↗Muse Spark 1.3
Output price / 1M$25 ↗$4.25 ↗Muse Spark 1.3
Cached input / 1M$0.5 ↗$0.15 ↗Muse Spark 1.3
Input + output / 1M
Lower is cheaper. List prices; batch, tool and regional fees excluded.
$30$5.5Muse Spark 1.3
Modalitiestext · visiontext · vision · videoMuse Spark 1.3
Released2025-11-242026-09-02—
Cited benchmark scores3929—
Reliability

Provider status

All providers →
More matchups

Claude Opus 4.5 vs …

More matchups

Muse Spark 1.3 vs …

Built by Respan
Which one wins on your data?

Public benchmarks are a starting point. Run Claude Opus 4.5 and Muse Spark 1.3 on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.