ModelsCompareBest forBenchmarksStatusPricingAPI
Compare / head-to-head

Claude Sonnet 4.5vsClaude Sonnet 4.6

Claude Sonnet 4.6 leads 15 of 17 shared benchmarks. Both list the same combined token price. Claude Sonnet 4.6 has the larger context window (1M tokens).

Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.

Shared benchmarks
2 – 15
Claude Sonnet 4.6 leads
Cheaper per token
Tie
list price, input + output
Larger context
Claude Sonnet 4.6
1M tokens
Provider uptime (30d)
100%
Anthropic
Claude Sonnet 4.5
Anthropic · released 2025-09-29
textvision
Context
200K
Max out
64K
Input /1M
$3
Output /1M
$15
Cached /1M
$0.3
Scores
21 · 10 core
Claude Sonnet 4.6
Anthropic · released 2026-02-17
textvision
Context
1M
Max out
128K
Input /1M
$3
Output /1M
$15
Cached /1M
$0.3
Scores
19 · 8 core
Quality

Benchmark matrix

BenchmarkClaude Sonnet 4.5Claude Sonnet 4.6ΔEdge
Reported by both · 17
SWE-bench Verified77.2% 79.6% -2.4 ptClaude Sonnet 4.6
GPQA Diamond83.4% 89.9% -6.5 ptClaude Sonnet 4.6
Humanity's Last Exam (no tools)17.7% 33.2% -15.5 ptClaude Sonnet 4.6
ARC-AGI-2 (Verified)13.6% 58.3% -44.7 ptClaude Sonnet 4.6
GDPval-AA1276 1633 -357Claude Sonnet 4.6
Humanity's Last Exam (with tools)33.6% 49% -15.4 ptClaude Sonnet 4.6
MCP Atlas43.8% 61.3% -17.5 ptClaude Sonnet 4.6
MMMLU89.5% 89.3% +0.2 ptClaude Sonnet 4.5
MMMU-Pro (no tools)63.4% 74.5% -11.1 ptClaude Sonnet 4.6
MMMU-Pro (with tools)68.9% 75.6% -6.7 ptClaude Sonnet 4.6
OSWorld-Verified61.4% 72.5% -11.1 ptClaude Sonnet 4.6
OTIS Mock AIME 2024-2025 (Epoch AI run)77.8% epoch85.8% epoch-8 ptClaude Sonnet 4.6
SimpleQA Verified30.7% epoch35.5% epoch-4.8 ptClaude Sonnet 4.6
SWE-bench Verified (Epoch AI run)71.3% epoch75.2% epoch-3.9 ptClaude Sonnet 4.6
Tau2-bench Retail86.2% 91.7% -5.5 ptClaude Sonnet 4.6
Tau2-bench Telecom98% 97.9% +0.1 ptClaude Sonnet 4.5
Terminal-Bench 2.051% 59.1% -8.1 ptClaude Sonnet 4.6
Only Claude Sonnet 4.5 reports · 3
AIME 202584.2% matharenanot reported
FrontierMath Tier 4 v2 (Epoch AI run)2.4% epochnot reported
FrontierMath Tiers 1-3 v2 (Epoch AI run)23.9% epochnot reported
Only Claude Sonnet 4.6 reports · 1
LMArena Elonot reported1458.3
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Claude Sonnet 4.5 minus Claude Sonnet 4.6 in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing

Side by side

SpecClaude Sonnet 4.5Claude Sonnet 4.6Edge
Context window200K tokens1M tokensClaude Sonnet 4.6
Max output64K tokens128K tokensClaude Sonnet 4.6
Input price / 1M$3 $3 Tie
Output price / 1M$15 $15 Tie
Cached input / 1M$0.3 $0.3 Tie
Input + output / 1M
Lower is cheaper. List prices; batch, tool and regional fees excluded.
$18$18Tie
Modalitiestext · visiontext · visionTie
Released2025-09-292026-02-17
Cited benchmark scores2119
Reliability

Anthropic status

All providers
More matchups

Claude Sonnet 4.5 vs …

More matchups

Claude Sonnet 4.6 vs …

Built by Respan
Which one wins on your data?

Public benchmarks are a starting point. Run Claude Sonnet 4.5 and Claude Sonnet 4.6 on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.