Compare / head-to-head

InklingvsMuse Spark 1.3

Muse Spark 1.3 leads 17 of 17 shared benchmarks. Inkling is 1.1x cheaper per token. Muse Spark 1.3 has the larger context window (1.0M tokens).

Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.

Shared benchmarks
0 – 17
Muse Spark 1.3 leads
Cheaper per token
Inkling
1.1x cheaper, input + output
Larger context
Muse Spark 1.3
1.0M tokens
Providers
2 providers
Thinking Machines Lab · Meta
Inkling
Thinking Machines Lab · released 2026-07-15
textvisionaudio
Context
1M
Max out
—
Input /1M
$1 ↗
Output /1M
$4.05 ↗
Cached /1M
$0.17
Scores
39 · 12 core
Muse Spark 1.3
Meta · released 2026-09-02
textvisionvideo
Context
1.0M
Max out
—
Input /1M
$1.25 ↗
Output /1M
$4.25 ↗
Cached /1M
$0.15
Scores
29 · 7 core
Quality

Benchmark matrix

BenchmarkInklingMuse Spark 1.3ΔEdge
Reported by both · 17
LMArena Elo1441.4 ↗1494.3 ↗-52.9Muse Spark 1.3
APEX-Agents (Mercor)33.8% ↗57.8% ↗-24 ptMuse Spark 1.3
Chess Puzzles (Epoch AI run)21% ↗epoch run38% ↗epoch run-17 ptMuse Spark 1.3
FrontierMath Tier 4 v2 (Epoch AI run)4.9% ↗epoch run46.3% ↗epoch run-41.4 ptMuse Spark 1.3
FrontierMath Tiers 1-3 v2 (Epoch AI run)33.3% ↗epoch run74.4% ↗epoch run-41.1 ptMuse Spark 1.3
LiveBench Agentic Coding (LiveBench)49.4% ↗64.1% ↗-14.7 ptMuse Spark 1.3
LiveBench Coding (LiveBench)71% ↗81.1% ↗-10.1 ptMuse Spark 1.3
LiveBench Data Analysis (LiveBench)72.8% ↗79.6% ↗-6.8 ptMuse Spark 1.3
LiveBench Instruction Following (LiveBench)70.1% ↗78% ↗-7.9 ptMuse Spark 1.3
LiveBench Language (LiveBench)73.5% ↗82.8% ↗-9.3 ptMuse Spark 1.3
LiveBench Mathematics (LiveBench)88.4% ↗95.9% ↗-7.5 ptMuse Spark 1.3
LiveBench Reasoning (LiveBench)78.3% ↗89.7% ↗-11.4 ptMuse Spark 1.3
LMArena Agent (LMArena)-0.1086 ↗0.04 ↗-0.1Muse Spark 1.3
LMArena WebDev (LMArena)1412.6 ↗1656.6 ↗-244Muse Spark 1.3
OTIS Mock AIME 2024-2025 (Epoch AI run)88.9% ↗epoch run99.2% ↗epoch run-10.3 ptMuse Spark 1.3
SimpleBench (SimpleBench)50% ↗81.8% ↗-31.8 ptMuse Spark 1.3
Terminal-Bench 4.0 (Vals AI)0.5% ↗24.7% ↗-24.2 ptMuse Spark 1.3
Only Inkling reports · 21
SWE-bench Verified77.6% ↗not reported——
GPQA Diamond88.3% ↗epoch runnot reported——
AIME 202697.1% ↗not reported——
Audio MC56.6% ↗not reported——
BrowseComp (w/ ctx management)77.1% ↗not reported——
CharXiv RQ78.1% ↗not reported——
CharXiv RQ (with python)82% ↗not reported——
FrontierSWE V2 (Proximal Labs)4.1% ↗not reported——
Global-MMLU-Lite88.7% ↗not reported——
IFBench79.8% ↗not reported——
MCP Atlas76% ↗not reported——
MMAU77.2% ↗not reported——
SAGE (Vals AI)36.6% ↗not reported——
SimpleQA Verified43.9% ↗not reported——
SWE-bench Pro (public)54.3% ↗not reported——
tau2-bench Banking Knowledge (Sierra)25% ↗not reported——
Terminal-Bench 2.1 (best harness)63.8% ↗not reported——
Toolathlon Verified45.5% ↗not reported——
Toolathlon-Verified (HKUST)45.5% ↗not reported——
VoiceBench91.4% ↗not reported——
WeirdML (Håvard Tveit Ihle)32.3% ↗not reported——
Only Muse Spark 1.3 reports · 12
AutomationBenchnot reported49.6% ↗——
DeepSearchQAnot reported90.3% ↗——
DeepSWE v1.1not reported75.4% ↗——
GDPval-AA v2not reported1754 ↗——
JobBenchnot reported64.9% ↗——
LMArena Vision (LMArena)not reported1290.2 ↗——
MRCR 256K-512Knot reported98.5% ↗——
MRCR 512K-1Mnot reported98.1% ↗——
Mystery Game Puzzles (Epoch AI run)not reported25% ↗epoch run——
OSWorld 2.0 (partial)not reported66.9% ↗——
SWE-Atlas Codebase QnAnot reported59.4% ↗——
Terminal-Bench 2.1not reported88.8% ↗——
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Inkling minus Muse Spark 1.3 in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing

Side by side

SpecInklingMuse Spark 1.3Edge
Context window1M tokens1.0M tokensMuse Spark 1.3
Max output———
Input price / 1M$1 ↗$1.25 ↗Inkling
Output price / 1M$4.05 ↗$4.25 ↗Inkling
Cached input / 1M$0.17 ↗$0.15 ↗Muse Spark 1.3
Input + output / 1M
Lower is cheaper. List prices; batch, tool and regional fees excluded.
$5.05$5.5Inkling
Modalitiestext · vision · audiotext · vision · videoTie
Released2026-07-152026-09-02—
Cited benchmark scores3929—
Reliability

Provider status

All providers →
More matchups

Inkling vs …

More matchups

Muse Spark 1.3 vs …

Built by Respan
Which one wins on your data?

Public benchmarks are a starting point. Run Inkling and Muse Spark 1.3 on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.