Compare / head-to-head

Muse Spark 1.1vsMuse Spark 1.3

Muse Spark 1.3 leads 15 of 15 shared benchmarks. Both list the same combined token price. Both offer a 1.0M-token context window.

Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.

Shared benchmarks
0 – 15
Muse Spark 1.3 leads
Cheaper per token
Tie
list price, input + output
Larger context
Tie
both 1.0M tokens
Providers
Meta
same provider
Muse Spark 1.1
Meta · released 2026-07-09
textvisionvideoaudio
Context
1.0M
Max out
—
Input /1M
$1.25 ↗
Output /1M
$4.25 ↗
Cached /1M
$0.15
Scores
30 · 7 core
Muse Spark 1.3
Meta · released 2026-09-02
textvisionvideo
Context
1.0M
Max out
—
Input /1M
$1.25 ↗
Output /1M
$4.25 ↗
Cached /1M
$0.15
Scores
29 · 7 core
Quality

Benchmark matrix

BenchmarkMuse Spark 1.1Muse Spark 1.3ΔEdge
Reported by both · 15
LMArena Elo1491.3 ↗1494.3 ↗-3Muse Spark 1.3
APEX-Agents (Mercor)31.8% ↗57.8% ↗-26 ptMuse Spark 1.3
DeepSWE v1.153.3% ↗75.4% ↗-22.1 ptMuse Spark 1.3
JobBench54.7% ↗64.9% ↗-10.2 ptMuse Spark 1.3
LiveBench Agentic Coding (LiveBench)58.5% ↗64.1% ↗-5.6 ptMuse Spark 1.3
LiveBench Coding (LiveBench)77.2% ↗81.1% ↗-3.9 ptMuse Spark 1.3
LiveBench Data Analysis (LiveBench)72.5% ↗79.6% ↗-7.1 ptMuse Spark 1.3
LiveBench Instruction Following (LiveBench)69.6% ↗78% ↗-8.4 ptMuse Spark 1.3
LiveBench Language (LiveBench)74.3% ↗82.8% ↗-8.5 ptMuse Spark 1.3
LiveBench Mathematics (LiveBench)87.1% ↗95.9% ↗-8.8 ptMuse Spark 1.3
LiveBench Reasoning (LiveBench)87.7% ↗89.7% ↗-2 ptMuse Spark 1.3
LMArena Agent (LMArena)-0.0488 ↗0.04 ↗-0.1Muse Spark 1.3
LMArena Vision (LMArena)1281.1 ↗1290.2 ↗-9.1Muse Spark 1.3
LMArena WebDev (LMArena)1541.7 ↗1656.6 ↗-114.9Muse Spark 1.3
Terminal-Bench 2.180% ↗88.8% ↗-8.8 ptMuse Spark 1.3
Only Muse Spark 1.1 reports · 15
BabyVision76.3% ↗not reported——
CharXiv Reasoning88.4% ↗not reported——
DeepSWE v1.1 (Datacurve)53.3% ↗not reported——
Finance Agent v257.2% ↗not reported——
Humanity's Last Exam (with tools)62.1% ↗not reported——
MCP Atlas88.1% ↗not reported——
OSWorld-Verified80.8% ↗not reported——
OSWorld-Verified (XLANG)80.7% ↗not reported——
SAGE (Vals AI)45.6% ↗not reported——
SimpleQA Verified57.8% ↗epoch runnot reported——
SWE-Bench Pro61.5% ↗not reported——
tau2-bench Banking Knowledge (Sierra)40.5% ↗not reported——
Toolathlon-Verified75.6% ↗not reported——
Toolathlon-Verified (HKUST)75.6% ↗not reported——
Vending-Bench 2 (Andon Labs)6520.48 ↗not reported——
Only Muse Spark 1.3 reports · 14
AutomationBenchnot reported49.6% ↗——
Chess Puzzles (Epoch AI run)not reported38% ↗epoch run——
DeepSearchQAnot reported90.3% ↗——
FrontierMath Tier 4 v2 (Epoch AI run)not reported46.3% ↗epoch run——
FrontierMath Tiers 1-3 v2 (Epoch AI run)not reported74.4% ↗epoch run——
GDPval-AA v2not reported1754 ↗——
MRCR 256K-512Knot reported98.5% ↗——
MRCR 512K-1Mnot reported98.1% ↗——
Mystery Game Puzzles (Epoch AI run)not reported25% ↗epoch run——
OSWorld 2.0 (partial)not reported66.9% ↗——
OTIS Mock AIME 2024-2025 (Epoch AI run)not reported99.2% ↗epoch run——
SimpleBench (SimpleBench)not reported81.8% ↗——
SWE-Atlas Codebase QnAnot reported59.4% ↗——
Terminal-Bench 4.0 (Vals AI)not reported24.7% ↗——
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Muse Spark 1.1 minus Muse Spark 1.3 in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing

Side by side

SpecMuse Spark 1.1Muse Spark 1.3Edge
Context window1.0M tokens1.0M tokensTie
Max output———
Input price / 1M$1.25 ↗$1.25 ↗Tie
Output price / 1M$4.25 ↗$4.25 ↗Tie
Cached input / 1M$0.15 ↗$0.15 ↗Tie
Input + output / 1M
Lower is cheaper. List prices; batch, tool and regional fees excluded.
$5.5$5.5Tie
Modalitiestext · vision · video · audiotext · vision · videoMuse Spark 1.1
Released2026-07-092026-09-02—
Cited benchmark scores3029—
Reliability

Meta status

All providers →
More matchups

Muse Spark 1.1 vs …

More matchups

Muse Spark 1.3 vs …

Built by Respan
Which one wins on your data?

Public benchmarks are a starting point. Run Muse Spark 1.1 and Muse Spark 1.3 on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.