Compare / head-to-head
Claude Opus 4.5vs
Claude Opus 4.6
Claude Opus 4.6 leads 22 of 30 shared benchmarks. Both list the same combined token price. Claude Opus 4.6 has the larger context window (1M tokens).
Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.
Shared benchmarks
7 – 22
Claude Opus 4.6 leads · 1 tied
Cheaper per token
Tie
list price, input + output
Larger context
Claude Opus 4.6
1M tokens
Providers
Anthropic
same provider
Claude Opus 4.5
Anthropic · released 2025-11-24
textvision
Claude Opus 4.6
Anthropic · released 2026-02-05
textvision
Quality
Benchmark matrix
30 shared · 8 only Claude Opus 4.5 · 8 only Claude Opus 4.6| Benchmark | Claude Opus 4.5 | Claude Opus 4.6 | Δ | Edge |
|---|---|---|---|---|
| Reported by both · 30 | ||||
| SWE-bench Verified | 80.9% ↗ | 80.8% ↗ | +0.1 pt | Claude Opus 4.5 |
| GPQA Diamond | 87% ↗ | 91.3% ↗ | -4.3 pt | Claude Opus 4.6 |
| LMArena Elo | 1473.6 ↗ | 1497.5 ↗ | -23.9 | Claude Opus 4.6 |
| ARC-AGI-2 (Verified) | 37.6% ↗ | 68.8% ↗ | -31.2 pt | Claude Opus 4.6 |
| Chess Puzzles (Epoch AI run) | 12% ↗epoch run | 17% ↗epoch run | -5 pt | Claude Opus 4.6 |
| EBR-bench (Epoch AI run) | 14.3% ↗epoch run | 12.7% ↗epoch run | +1.6 pt | Claude Opus 4.5 |
| FrontierMath Tier 4 v2 (Epoch AI run) | 4.9% ↗epoch run | 26.8% ↗epoch run | -21.9 pt | Claude Opus 4.6 |
| FrontierMath Tiers 1-3 v2 (Epoch AI run) | 34.4% ↗epoch run | 66% ↗epoch run | -31.6 pt | Claude Opus 4.6 |
| Furniture Assembly (Epoch AI run) | 28.3% ↗epoch run | 28.3% ↗epoch run | 0 | Tie |
| GSO Opt@1 (GSO) | 24.5% ↗ | 37.3% ↗ | -12.8 pt | Claude Opus 4.6 |
| LiveBench Agentic Coding (LiveBench) | 39.7% ↗ | 49% ↗ | -9.3 pt | Claude Opus 4.6 |
| LiveBench Coding (LiveBench) | 79.7% ↗ | 78.2% ↗ | +1.5 pt | Claude Opus 4.5 |
| LiveBench Data Analysis (LiveBench) | 74.4% ↗ | 69.9% ↗ | +4.5 pt | Claude Opus 4.5 |
| LiveBench Instruction Following (LiveBench) | 62.5% ↗ | 63.3% ↗ | -0.8 pt | Claude Opus 4.6 |
| LiveBench Language (LiveBench) | 81.3% ↗ | 83.3% ↗ | -2 pt | Claude Opus 4.6 |
| LiveBench Mathematics (LiveBench) | 90.4% ↗ | 89.3% ↗ | +1.1 pt | Claude Opus 4.5 |
| LiveBench Reasoning (LiveBench) | 80.1% ↗ | 88.7% ↗ | -8.6 pt | Claude Opus 4.6 |
| LMArena WebDev (LMArena) | 1493.3 ↗ | 1545.9 ↗ | -52.6 | Claude Opus 4.6 |
| MCP Atlas | 62.3% ↗ | 59.5% ↗ | +2.8 pt | Claude Opus 4.5 |
| MMMLU | 90.8% ↗ | 91.1% ↗ | -0.3 pt | Claude Opus 4.6 |
| Mystery Game Puzzles (Epoch AI run) | 22% ↗epoch run | 25% ↗epoch run | -3 pt | Claude Opus 4.6 |
| OTIS Mock AIME 2024-2025 (Epoch AI run) | 86.1% ↗epoch run | 94.4% ↗epoch run | -8.3 pt | Claude Opus 4.6 |
| SAGE (Vals AI) | 52.1% ↗ | 51.6% ↗ | +0.5 pt | Claude Opus 4.5 |
| SimpleBench (SimpleBench) | 62% ↗ | 67.6% ↗ | -5.6 pt | Claude Opus 4.6 |
| SimpleQA Verified | 45.7% ↗epoch run | 47% ↗epoch run | -1.3 pt | Claude Opus 4.6 |
| SWE-bench Verified (Epoch AI run) | 76.7% ↗epoch run | 78.7% ↗epoch run | -2 pt | Claude Opus 4.6 |
| tau2-bench Banking Knowledge (Sierra) | 24.7% ↗ | 27.3% ↗ | -2.6 pt | Claude Opus 4.6 |
| Terminal-Bench 2.0 | 59.3% ↗ | 65.4% ↗ | -6.1 pt | Claude Opus 4.6 |
| Vending-Bench 2 (Andon Labs) | 4967.06 ↗ | 8017.59 ↗ | -3050.5 | Claude Opus 4.6 |
| WeirdML (Håvard Tveit Ihle) | 63.7% ↗ | 78% ↗ | -14.3 pt | Claude Opus 4.6 |
| Only Claude Opus 4.5 reports · 8 | ||||
| BALROG (BALROG) | 43.5% ↗ | not reported | — | — |
| MMMU (validation) | 80.7% ↗ | not reported | — | — |
| OSWorld | 66.3% ↗ | not reported | — | — |
| tau2-bench Airline (Sierra) | 84% ↗ | not reported | — | — |
| tau2-bench Retail (Sierra) | 79.6% ↗ | not reported | — | — |
| tau2-bench Telecom (Sierra) | 92.3% ↗ | not reported | — | — |
| τ2-bench (Retail) | 88.9% ↗ | not reported | — | — |
| τ2-bench (Telecom) | 98.2% ↗ | not reported | — | — |
| Only Claude Opus 4.6 reports · 8 | ||||
| AIME 2026 | not reported | 96.7% ↗matharena | — | — |
| APEX-Agents (Mercor) | not reported | 46.3% ↗ | — | — |
| LMArena Vision (LMArena) | not reported | 1299.4 ↗ | — | — |
| MMMU-Pro (no tools) | not reported | 73.9% ↗ | — | — |
| MMMU-Pro (with tools) | not reported | 77.3% ↗ | — | — |
| OSWorld-Verified | not reported | 72.7% ↗ | — | — |
| Tau2-bench Retail | not reported | 91.9% ↗ | — | — |
| Tau2-bench Telecom | not reported | 99.3% ↗ | — | — |
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Claude Opus 4.5 minus Claude Opus 4.6 in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing
Side by side
Official pricing pages and model cards| Spec | Claude Opus 4.5 | Claude Opus 4.6 | Edge |
|---|---|---|---|
| Context window | 200K tokens | 1M tokens | Claude Opus 4.6 |
| Max output | 64K tokens | 128K tokens | Claude Opus 4.6 |
| Input price / 1M | $5 ↗ | $5 ↗ | Tie |
| Output price / 1M | $25 ↗ | $25 ↗ | Tie |
| Cached input / 1M | $0.5 ↗ | $0.5 ↗ | Tie |
| Input + output / 1M Lower is cheaper. List prices; batch, tool and regional fees excluded. | $30 | $30 | Tie |
| Modalities | text · vision | text · vision | Tie |
| Released | 2025-11-24 | 2026-02-05 | — |
| Cited benchmark scores | 39 | 46 | — |
Reliability
Anthropic status
Live from /statusMore matchups
Claude Opus 4.5 vs …
Models sharing the most benchmarksMore matchups
Claude Opus 4.6 vs …
Models sharing the most benchmarksBuilt by Respan
Which one wins on your data?
Public benchmarks are a starting point. Run Claude Opus 4.5 and Claude Opus 4.6 on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.