Compare / head-to-head

Claude Opus 4.6vsClaude Opus 4.8

Claude Opus 4.8 leads 23 of 29 shared benchmarks. Both list the same combined token price. Both offer a 1M-token context window.

Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.

Shared benchmarks
6 – 23
Claude Opus 4.8 leads
Cheaper per token
Tie
list price, input + output
Larger context
Tie
both 1M tokens
Providers
Anthropic
same provider
Claude Opus 4.6
Anthropic · released 2026-02-05
textvision
Context
1M
Max out
128K
Input /1M
$5 ↗
Output /1M
$25 ↗
Cached /1M
$0.5
Scores
46 · 11 core
Claude Opus 4.8
Anthropic · released 2026-05-28
textvision
Context
1M
Max out
128K
Input /1M
$5 ↗
Output /1M
$25 ↗
Cached /1M
$0.5
Scores
47 · 16 core
Quality

Benchmark matrix

BenchmarkClaude Opus 4.6Claude Opus 4.8ΔEdge
Reported by both · 29
SWE-bench Verified80.8% ↗88.6% ↗-7.8 ptClaude Opus 4.8
GPQA Diamond91.3% ↗93.6% ↗-2.3 ptClaude Opus 4.8
AIME 202696.7% ↗matharena100% ↗matharena ⚠-3.3 ptClaude Opus 4.8
LMArena Elo1497.5 ↗1452.7 ↗+44.8Claude Opus 4.6
APEX-Agents (Mercor)46.3% ↗48.9% ↗-2.6 ptClaude Opus 4.8
Chess Puzzles (Epoch AI run)17% ↗epoch run34% ↗epoch run-17 ptClaude Opus 4.8
EBR-bench (Epoch AI run)12.7% ↗epoch run28.6% ↗epoch run-15.9 ptClaude Opus 4.8
FrontierMath Tier 4 v2 (Epoch AI run)26.8% ↗epoch run56.1% ↗epoch run-29.3 ptClaude Opus 4.8
FrontierMath Tiers 1-3 v2 (Epoch AI run)66% ↗epoch run80% ↗epoch run-14 ptClaude Opus 4.8
Furniture Assembly (Epoch AI run)28.3% ↗epoch run42.5% ↗epoch run-14.2 ptClaude Opus 4.8
GSO Opt@1 (GSO)37.3% ↗47.1% ↗-9.8 ptClaude Opus 4.8
LiveBench Agentic Coding (LiveBench)49% ↗50.5% ↗-1.5 ptClaude Opus 4.8
LiveBench Coding (LiveBench)78.2% ↗81.8% ↗-3.6 ptClaude Opus 4.8
LiveBench Data Analysis (LiveBench)69.9% ↗66% ↗+3.9 ptClaude Opus 4.6
LiveBench Instruction Following (LiveBench)63.3% ↗72% ↗-8.7 ptClaude Opus 4.8
LiveBench Language (LiveBench)83.3% ↗79.7% ↗+3.6 ptClaude Opus 4.6
LiveBench Mathematics (LiveBench)89.3% ↗94.3% ↗-5 ptClaude Opus 4.8
LiveBench Reasoning (LiveBench)88.7% ↗89.2% ↗-0.5 ptClaude Opus 4.8
LMArena Vision (LMArena)1299.4 ↗1286.4 ↗+13Claude Opus 4.6
LMArena WebDev (LMArena)1545.9 ↗1555.5 ↗-9.6Claude Opus 4.8
Mystery Game Puzzles (Epoch AI run)25% ↗epoch run36% ↗epoch run-11 ptClaude Opus 4.8
OSWorld-Verified72.7% ↗83.4% ↗-10.7 ptClaude Opus 4.8
OTIS Mock AIME 2024-2025 (Epoch AI run)94.4% ↗epoch run98.3% ↗epoch run-3.9 ptClaude Opus 4.8
SAGE (Vals AI)51.6% ↗54.8% ↗-3.2 ptClaude Opus 4.8
SimpleBench (SimpleBench)67.6% ↗64.8% ↗+2.8 ptClaude Opus 4.6
SimpleQA Verified47% ↗epoch run53% ↗epoch run-6 ptClaude Opus 4.8
tau2-bench Banking Knowledge (Sierra)27.3% ↗39.7% ↗-12.4 ptClaude Opus 4.8
Vending-Bench 2 (Andon Labs)8017.59 ↗5787.43 ↗+2230.2Claude Opus 4.6
WeirdML (Håvard Tveit Ihle)78% ↗82.9% ↗-4.9 ptClaude Opus 4.8
Only Claude Opus 4.6 reports · 9
ARC-AGI-2 (Verified)68.8% ↗not reported——
MCP Atlas59.5% ↗not reported——
MMMLU91.1% ↗not reported——
MMMU-Pro (no tools)73.9% ↗not reported——
MMMU-Pro (with tools)77.3% ↗not reported——
SWE-bench Verified (Epoch AI run)78.7% ↗epoch runnot reported——
Tau2-bench Retail91.9% ↗not reported——
Tau2-bench Telecom99.3% ↗not reported——
Terminal-Bench 2.065.4% ↗not reported——
Only Claude Opus 4.8 reports · 11
Humanity's Last Exam (no tools)not reported49.8% ↗——
BrowseCompnot reported84.3% ↗——
DeepSWE v1.1 (Datacurve)not reported59% ↗——
Humanity's Last Exam (with tools)not reported57.9% ↗——
LMArena Agent (LMArena)not reported0.0664 ↗——
SWE-bench Multilingualnot reported84.4% ↗——
SWE-bench Multimodalnot reported38.4% ↗——
SWE-Bench Pronot reported69.2% ↗——
Terminal-Bench 2.1not reported74.6% ↗——
Terminal-Bench 4.0 (Vals AI)not reported23.2% ↗——
Toolathlon-Verified (HKUST)not reported76.2% ↗——
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Claude Opus 4.6 minus Claude Opus 4.8 in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing

Side by side

SpecClaude Opus 4.6Claude Opus 4.8Edge
Context window1M tokens1M tokensTie
Max output128K tokens128K tokensTie
Input price / 1M$5 ↗$5 ↗Tie
Output price / 1M$25 ↗$25 ↗Tie
Cached input / 1M$0.5 ↗$0.5 ↗Tie
Input + output / 1M
Lower is cheaper. List prices; batch, tool and regional fees excluded.
$30$30Tie
Modalitiestext · visiontext · visionTie
Released2026-02-052026-05-28—
Cited benchmark scores4647—
Reliability

Anthropic status

All providers →
More matchups

Claude Opus 4.6 vs …

More matchups

Claude Opus 4.8 vs …

Built by Respan
Which one wins on your data?

Public benchmarks are a starting point. Run Claude Opus 4.6 and Claude Opus 4.8 on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.