Compare / head-to-head

Claude Opus 4.5vsClaude Sonnet 5.5

Claude Sonnet 5.5 leads 16 of 18 shared benchmarks. Claude Sonnet 5.5 is 2.5x cheaper per token. Claude Sonnet 5.5 has the larger context window (1M tokens).

Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.

Shared benchmarks
2 – 16
Claude Sonnet 5.5 leads
Cheaper per token
Claude Sonnet 5.5
2.5x cheaper, input + output
Larger context
Claude Sonnet 5.5
1M tokens
Providers
Anthropic
same provider
Claude Opus 4.5
Anthropic · released 2025-11-24
textvision
Context
200K
Max out
64K
Input /1M
$5 ↗
Output /1M
$25 ↗
Cached /1M
$0.5
Scores
39 · 9 core
Claude Sonnet 5.5
Anthropic · released 2026-09-28
textvision
Context
1M
Max out
128K
Input /1M
$2 ↗
Output /1M
$10 ↗
Cached /1M
$0.2
Scores
29 · 10 core
Quality

Benchmark matrix

BenchmarkClaude Opus 4.5Claude Sonnet 5.5ΔEdge
Reported by both · 18
GPQA Diamond87% ↗95.6% ↗epoch run-8.6 ptClaude Sonnet 5.5
LMArena Elo1473.6 ↗1471 ↗+2.6Claude Opus 4.5
FrontierMath Tier 4 v2 (Epoch AI run)4.9% ↗epoch run80.5% ↗epoch run-75.6 ptClaude Sonnet 5.5
FrontierMath Tiers 1-3 v2 (Epoch AI run)34.4% ↗epoch run88.8% ↗epoch run-54.4 ptClaude Sonnet 5.5
Furniture Assembly (Epoch AI run)28.3% ↗epoch run75% ↗epoch run-46.7 ptClaude Sonnet 5.5
LiveBench Agentic Coding (LiveBench)39.7% ↗56.3% ↗-16.6 ptClaude Sonnet 5.5
LiveBench Coding (LiveBench)79.7% ↗91.4% ↗-11.7 ptClaude Sonnet 5.5
LiveBench Data Analysis (LiveBench)74.4% ↗78.6% ↗-4.2 ptClaude Sonnet 5.5
LiveBench Instruction Following (LiveBench)62.5% ↗70.5% ↗-8 ptClaude Sonnet 5.5
LiveBench Language (LiveBench)81.3% ↗83.4% ↗-2.1 ptClaude Sonnet 5.5
LiveBench Mathematics (LiveBench)90.4% ↗96.7% ↗-6.3 ptClaude Sonnet 5.5
LiveBench Reasoning (LiveBench)80.1% ↗91.6% ↗-11.5 ptClaude Sonnet 5.5
LMArena WebDev (LMArena)1493.3 ↗1786.3 ↗-293Claude Sonnet 5.5
Mystery Game Puzzles (Epoch AI run)22% ↗epoch run65% ↗epoch run-43 ptClaude Sonnet 5.5
OTIS Mock AIME 2024-2025 (Epoch AI run)86.1% ↗epoch run100% ↗epoch run-13.9 ptClaude Sonnet 5.5
SAGE (Vals AI)52.1% ↗51.8% ↗+0.3 ptClaude Opus 4.5
SimpleBench (SimpleBench)62% ↗75.9% ↗-13.9 ptClaude Sonnet 5.5
SimpleQA Verified45.7% ↗epoch run46.5% ↗epoch run-0.8 ptClaude Sonnet 5.5
Only Claude Opus 4.5 reports · 20
SWE-bench Verified80.9% ↗not reported——
ARC-AGI-2 (Verified)37.6% ↗not reported——
BALROG (BALROG)43.5% ↗not reported——
Chess Puzzles (Epoch AI run)12% ↗epoch runnot reported——
EBR-bench (Epoch AI run)14.3% ↗epoch runnot reported——
GSO Opt@1 (GSO)24.5% ↗not reported——
MCP Atlas62.3% ↗not reported——
MMMLU90.8% ↗not reported——
MMMU (validation)80.7% ↗not reported——
OSWorld66.3% ↗not reported——
SWE-bench Verified (Epoch AI run)76.7% ↗epoch runnot reported——
tau2-bench Airline (Sierra)84% ↗not reported——
tau2-bench Banking Knowledge (Sierra)24.7% ↗not reported——
tau2-bench Retail (Sierra)79.6% ↗not reported——
tau2-bench Telecom (Sierra)92.3% ↗not reported——
Terminal-Bench 2.059.3% ↗not reported——
Vending-Bench 2 (Andon Labs)4967.06 ↗not reported——
WeirdML (Håvard Tveit Ihle)63.7% ↗not reported——
τ2-bench (Retail)88.9% ↗not reported——
τ2-bench (Telecom)98.2% ↗not reported——
Only Claude Sonnet 5.5 reports · 11
APEX-Agents (Mercor)not reported75.5% ↗——
Chartography (no tools)not reported61.6% ↗——
CursorBench 4.0not reported55.5% ↗——
FrontierCode v1.1 (Main)not reported46.2% ↗——
FrontierSWE V2 (Proximal Labs)not reported61.9% ↗——
Humanity's Last Exam (with tools)not reported64.5% ↗——
LMArena Agent (LMArena)not reported0.1252 ↗——
LMArena Vision (LMArena)not reported1268.3 ↗——
OSWorld 2.1 (partial)not reported80.1% ↗——
Terminal-Bench 4.0not reported70.6% ↗——
Terminal-Bench 4.0 (Vals AI)not reported64.1% ↗——
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Claude Opus 4.5 minus Claude Sonnet 5.5 in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing

Side by side

SpecClaude Opus 4.5Claude Sonnet 5.5Edge
Context window200K tokens1M tokensClaude Sonnet 5.5
Max output64K tokens128K tokensClaude Sonnet 5.5
Input price / 1M$5 ↗$2 ↗Claude Sonnet 5.5
Output price / 1M$25 ↗$10 ↗Claude Sonnet 5.5
Cached input / 1M$0.5 ↗$0.2 ↗Claude Sonnet 5.5
Input + output / 1M
Lower is cheaper. List prices; batch, tool and regional fees excluded.
$30$12Claude Sonnet 5.5
Modalitiestext · visiontext · visionTie
Released2025-11-242026-09-28—
Cited benchmark scores3929—
Reliability

Anthropic status

All providers →
More matchups

Claude Opus 4.5 vs …

More matchups

Claude Sonnet 5.5 vs …

Built by Respan
Which one wins on your data?

Public benchmarks are a starting point. Run Claude Opus 4.5 and Claude Sonnet 5.5 on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.