Compare / head-to-head

Claude Opus 4.8vsGemini 2.5 Pro

Claude Opus 4.8 leads 14 of 15 shared benchmarks. Gemini 2.5 Pro is 2.7x cheaper per token. Gemini 2.5 Pro has the larger context window (1.0M tokens).

Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.

Shared benchmarks
14 – 1
Claude Opus 4.8 leads
Cheaper per token
Gemini 2.5 Pro
2.7x cheaper, input + output
Larger context
Gemini 2.5 Pro
1.0M tokens
Providers
2 providers
Anthropic · Google
Claude Opus 4.8
Anthropic · released 2026-05-28
textvision
Context
1M
Max out
128K
Input /1M
$5 ↗
Output /1M
$25 ↗
Cached /1M
$0.5
Scores
47 · 16 core
Gemini 2.5 Pro
Google · released 2025-06-17
textvisionaudio
Context
1.0M
Max out
66K
Input /1M
$1.25 ↗
Output /1M
$10 ↗
Cached /1M
$0.125
Scores
28 · 11 core
Quality

Benchmark matrix

BenchmarkClaude Opus 4.8Gemini 2.5 ProΔEdge
Reported by both · 15
SWE-bench Verified88.6% ↗59.6% ↗+29 ptClaude Opus 4.8
GPQA Diamond93.6% ↗86.4% ↗+7.2 ptClaude Opus 4.8
Humanity's Last Exam (no tools)49.8% ↗21.6% ↗+28.2 ptClaude Opus 4.8
LMArena Elo1452.7 ↗1457.8 ↗-5.1Gemini 2.5 Pro
Chess Puzzles (Epoch AI run)34% ↗epoch run20% ↗epoch run+14 ptClaude Opus 4.8
FrontierMath Tier 4 v2 (Epoch AI run)56.1% ↗epoch run0% ↗epoch run+56.1 ptClaude Opus 4.8
FrontierMath Tiers 1-3 v2 (Epoch AI run)80% ↗epoch run24.6% ↗epoch run+55.4 ptClaude Opus 4.8
GSO Opt@1 (GSO)47.1% ↗0% ↗+47.1 ptClaude Opus 4.8
LMArena Vision (LMArena)1286.4 ↗1247.5 ↗+38.9Claude Opus 4.8
LMArena WebDev (LMArena)1555.5 ↗1227 ↗+328.5Claude Opus 4.8
OTIS Mock AIME 2024-2025 (Epoch AI run)98.3% ↗epoch run84.7% ↗epoch run+13.6 ptClaude Opus 4.8
SAGE (Vals AI)54.8% ↗41.9% ↗+12.9 ptClaude Opus 4.8
tau2-bench Banking Knowledge (Sierra)39.7% ↗13.7% ↗+26 ptClaude Opus 4.8
Vending-Bench 2 (Andon Labs)5787.43 ↗573.64 ↗+5213.8Claude Opus 4.8
WeirdML (Håvard Tveit Ihle)82.9% ↗54% ↗+28.9 ptClaude Opus 4.8
Only Claude Opus 4.8 reports · 25
AIME 2026100% ↗matharena ⚠not reported——
APEX-Agents (Mercor)48.9% ↗not reported——
BrowseComp84.3% ↗not reported——
DeepSWE v1.1 (Datacurve)59% ↗not reported——
EBR-bench (Epoch AI run)28.6% ↗epoch runnot reported——
Furniture Assembly (Epoch AI run)42.5% ↗epoch runnot reported——
Humanity's Last Exam (with tools)57.9% ↗not reported——
LiveBench Agentic Coding (LiveBench)50.5% ↗not reported——
LiveBench Coding (LiveBench)81.8% ↗not reported——
LiveBench Data Analysis (LiveBench)66% ↗not reported——
LiveBench Instruction Following (LiveBench)72% ↗not reported——
LiveBench Language (LiveBench)79.7% ↗not reported——
LiveBench Mathematics (LiveBench)94.3% ↗not reported——
LiveBench Reasoning (LiveBench)89.2% ↗not reported——
LMArena Agent (LMArena)0.0664 ↗not reported——
Mystery Game Puzzles (Epoch AI run)36% ↗epoch runnot reported——
OSWorld-Verified83.4% ↗not reported——
SimpleBench (SimpleBench)64.8% ↗not reported——
SimpleQA Verified53% ↗epoch runnot reported——
SWE-bench Multilingual84.4% ↗not reported——
SWE-bench Multimodal38.4% ↗not reported——
SWE-Bench Pro69.2% ↗not reported——
Terminal-Bench 2.174.6% ↗not reported——
Terminal-Bench 4.0 (Vals AI)23.2% ↗not reported——
Toolathlon-Verified (HKUST)76.2% ↗not reported——
Only Gemini 2.5 Pro reports · 5
AIME 2025not reported88% ↗——
Aider Polyglotnot reported82.2% ↗——
LiveCodeBenchnot reported74.2% ↗——
MMMUnot reported82% ↗——
SWE-bench Verified (Epoch AI run)not reported57.6% ↗epoch run——
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Claude Opus 4.8 minus Gemini 2.5 Pro in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing

Side by side

SpecClaude Opus 4.8Gemini 2.5 ProEdge
Context window1M tokens1.0M tokensGemini 2.5 Pro
Max output128K tokens66K tokensClaude Opus 4.8
Input price / 1M$5 ↗$1.25 ↗Gemini 2.5 Pro
Output price / 1M$25 ↗$10 ↗Gemini 2.5 Pro
Cached input / 1M$0.5 ↗$0.125 ↗Gemini 2.5 Pro
Input + output / 1M
Lower is cheaper. List prices; batch, tool and regional fees excluded.
$30$11.25Gemini 2.5 Pro
Modalitiestext · visiontext · vision · audioGemini 2.5 Pro
Released2026-05-282025-06-17—
Cited benchmark scores4728—
Reliability

Provider status

All providers →
More matchups

Claude Opus 4.8 vs …

More matchups

Gemini 2.5 Pro vs …

Built by Respan
Which one wins on your data?

Public benchmarks are a starting point. Run Claude Opus 4.8 and Gemini 2.5 Pro on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.