Compare / head-to-head

Claude Opus 4.7vsMuse Spark 1.1

Claude Opus 4.7 leads 11 of 19 shared benchmarks. Muse Spark 1.1 is 5.5x cheaper per token. Muse Spark 1.1 has the larger context window (1.0M tokens).

Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.

Shared benchmarks
11 – 8
Claude Opus 4.7 leads
Cheaper per token
Muse Spark 1.1
5.5x cheaper, input + output
Larger context
Muse Spark 1.1
1.0M tokens
Providers
2 providers
Anthropic · Meta
Claude Opus 4.7
Anthropic · released 2026-04-16
textvision
Context
1M
Max out
128K
Input /1M
$5 ↗
Output /1M
$25 ↗
Cached /1M
$0.5
Scores
46 · 16 core
Muse Spark 1.1
Meta · released 2026-07-09
textvisionvideoaudio
Context
1.0M
Max out
—
Input /1M
$1.25 ↗
Output /1M
$4.25 ↗
Cached /1M
$0.15
Scores
30 · 7 core
Quality

Benchmark matrix

BenchmarkClaude Opus 4.7Muse Spark 1.1ΔEdge
Reported by both · 19
LMArena Elo1483.4 ↗1491.3 ↗-7.9Muse Spark 1.1
APEX-Agents (Mercor)49.2% ↗31.8% ↗+17.4 ptClaude Opus 4.7
Humanity's Last Exam (with tools)54.7% ↗62.1% ↗-7.4 ptMuse Spark 1.1
LiveBench Agentic Coding (LiveBench)50.7% ↗58.5% ↗-7.8 ptMuse Spark 1.1
LiveBench Coding (LiveBench)82.1% ↗77.2% ↗+4.9 ptClaude Opus 4.7
LiveBench Data Analysis (LiveBench)78.3% ↗72.5% ↗+5.8 ptClaude Opus 4.7
LiveBench Instruction Following (LiveBench)66.7% ↗69.6% ↗-2.9 ptMuse Spark 1.1
LiveBench Language (LiveBench)77.9% ↗74.3% ↗+3.6 ptClaude Opus 4.7
LiveBench Mathematics (LiveBench)92.9% ↗87.1% ↗+5.8 ptClaude Opus 4.7
LiveBench Reasoning (LiveBench)87.2% ↗87.7% ↗-0.5 ptMuse Spark 1.1
LMArena Vision (LMArena)1298.1 ↗1281.1 ↗+17Claude Opus 4.7
LMArena WebDev (LMArena)1557.3 ↗1541.7 ↗+15.6Claude Opus 4.7
OSWorld-Verified82.8% ↗80.8% ↗+2 ptClaude Opus 4.7
SAGE (Vals AI)56.1% ↗45.6% ↗+10.5 ptClaude Opus 4.7
SimpleQA Verified51.7% ↗epoch run57.8% ↗epoch run-6.1 ptMuse Spark 1.1
SWE-Bench Pro64.3% ↗61.5% ↗+2.8 ptClaude Opus 4.7
tau2-bench Banking Knowledge (Sierra)40.2% ↗40.5% ↗-0.3 ptMuse Spark 1.1
Terminal-Bench 2.166.1% ↗80% ↗-13.9 ptMuse Spark 1.1
Vending-Bench 2 (Andon Labs)10936.76 ↗6520.48 ↗+4416.3Claude Opus 4.7
Only Claude Opus 4.7 reports · 19
SWE-bench Verified87.6% ↗not reported——
GPQA Diamond94.2% ↗not reported——
AIME 202695.8% ↗matharena ⚠not reported——
Humanity's Last Exam (no tools)46.9% ↗not reported——
BrowseComp79.8% ↗not reported——
Chess Puzzles (Epoch AI run)30% ↗epoch runnot reported——
EBR-bench (Epoch AI run)19% ↗epoch runnot reported——
FrontierMath Tier 4 v2 (Epoch AI run)31.7% ↗epoch runnot reported——
FrontierMath Tiers 1-3 v2 (Epoch AI run)70.2% ↗epoch runnot reported——
Furniture Assembly (Epoch AI run)33.3% ↗epoch runnot reported——
GSO Opt@1 (GSO)42.2% ↗not reported——
MirrorCode (Epoch AI run)31.1% ↗epoch runnot reported——
Mystery Game Puzzles (Epoch AI run)28% ↗epoch runnot reported——
OTIS Mock AIME 2024-2025 (Epoch AI run)97.8% ↗epoch runnot reported——
SimpleBench (SimpleBench)61.7% ↗not reported——
SWE-bench Multilingual80.5% ↗not reported——
SWE-bench Multimodal34.5% ↗not reported——
SWE-bench Verified (Epoch AI run)83.5% ↗epoch runnot reported——
WeirdML (Håvard Tveit Ihle)76.4% ↗not reported——
Only Muse Spark 1.1 reports · 11
BabyVisionnot reported76.3% ↗——
CharXiv Reasoningnot reported88.4% ↗——
DeepSWE v1.1not reported53.3% ↗——
DeepSWE v1.1 (Datacurve)not reported53.3% ↗——
Finance Agent v2not reported57.2% ↗——
JobBenchnot reported54.7% ↗——
LMArena Agent (LMArena)not reported-0.0488 ↗——
MCP Atlasnot reported88.1% ↗——
OSWorld-Verified (XLANG)not reported80.7% ↗——
Toolathlon-Verifiednot reported75.6% ↗——
Toolathlon-Verified (HKUST)not reported75.6% ↗——
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Claude Opus 4.7 minus Muse Spark 1.1 in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing

Side by side

SpecClaude Opus 4.7Muse Spark 1.1Edge
Context window1M tokens1.0M tokensMuse Spark 1.1
Max output128K tokens——
Input price / 1M$5 ↗$1.25 ↗Muse Spark 1.1
Output price / 1M$25 ↗$4.25 ↗Muse Spark 1.1
Cached input / 1M$0.5 ↗$0.15 ↗Muse Spark 1.1
Input + output / 1M
Lower is cheaper. List prices; batch, tool and regional fees excluded.
$30$5.5Muse Spark 1.1
Modalitiestext · visiontext · vision · video · audioMuse Spark 1.1
Released2026-04-162026-07-09—
Cited benchmark scores4630—
Reliability

Provider status

All providers →
More matchups

Claude Opus 4.7 vs …

More matchups

Muse Spark 1.1 vs …

Built by Respan
Which one wins on your data?

Public benchmarks are a starting point. Run Claude Opus 4.7 and Muse Spark 1.1 on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.