Compare / head-to-head

Claude Opus 4.8vsGemini 3.5 Flash

Claude Opus 4.8 leads 22 of 28 shared benchmarks. Gemini 3.5 Flash is 2.9x cheaper per token. Gemini 3.5 Flash has the larger context window (1.0M tokens).

Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.

Shared benchmarks
22 – 6
Claude Opus 4.8 leads
Cheaper per token
Gemini 3.5 Flash
2.9x cheaper, input + output
Larger context
Gemini 3.5 Flash
1.0M tokens
Providers
2 providers
Anthropic · Google
Claude Opus 4.8
Anthropic · released 2026-05-28
textvision
Context
1M
Max out
128K
Input /1M
$5 ↗
Output /1M
$25 ↗
Cached /1M
$0.5
Scores
47 · 16 core
Gemini 3.5 Flash
Google · released 2026-05-19
textvisionaudio
Context
1.0M
Max out
66K
Input /1M
$1.5 ↗
Output /1M
$9 ↗
Cached /1M
$0.15
Scores
42 · 12 core
Quality

Benchmark matrix

BenchmarkClaude Opus 4.8Gemini 3.5 FlashΔEdge
Reported by both · 28
GPQA Diamond93.6% ↗92.8% ↗epoch run+0.8 ptClaude Opus 4.8
AIME 2026100% ↗matharena ⚠95% ↗matharena ⚠+5 ptClaude Opus 4.8
LMArena Elo1452.7 ↗1477.4 ↗-24.7Gemini 3.5 Flash
APEX-Agents (Mercor)48.9% ↗27.5% ↗+21.4 ptClaude Opus 4.8
Chess Puzzles (Epoch AI run)34% ↗epoch run50% ↗epoch run-16 ptGemini 3.5 Flash
DeepSWE v1.1 (Datacurve)59% ↗36.1% ↗+22.9 ptClaude Opus 4.8
EBR-bench (Epoch AI run)28.6% ↗epoch run4.8% ↗epoch run+23.8 ptClaude Opus 4.8
FrontierMath Tier 4 v2 (Epoch AI run)56.1% ↗epoch run26.8% ↗epoch run+29.3 ptClaude Opus 4.8
FrontierMath Tiers 1-3 v2 (Epoch AI run)80% ↗epoch run62.8% ↗epoch run+17.2 ptClaude Opus 4.8
LiveBench Agentic Coding (LiveBench)50.5% ↗49% ↗+1.5 ptClaude Opus 4.8
LiveBench Coding (LiveBench)81.8% ↗78.2% ↗+3.6 ptClaude Opus 4.8
LiveBench Data Analysis (LiveBench)66% ↗64.9% ↗+1.1 ptClaude Opus 4.8
LiveBench Instruction Following (LiveBench)72% ↗75.6% ↗-3.6 ptGemini 3.5 Flash
LiveBench Language (LiveBench)79.7% ↗84.6% ↗-4.9 ptGemini 3.5 Flash
LiveBench Mathematics (LiveBench)94.3% ↗88.2% ↗+6.1 ptClaude Opus 4.8
LiveBench Reasoning (LiveBench)89.2% ↗82% ↗+7.2 ptClaude Opus 4.8
LMArena Vision (LMArena)1286.4 ↗1284.4 ↗+2Claude Opus 4.8
LMArena WebDev (LMArena)1555.5 ↗1499 ↗+56.5Claude Opus 4.8
Mystery Game Puzzles (Epoch AI run)36% ↗epoch run32% ↗epoch run+4 ptClaude Opus 4.8
OSWorld-Verified83.4% ↗78.4% ↗+5 ptClaude Opus 4.8
OTIS Mock AIME 2024-2025 (Epoch AI run)98.3% ↗epoch run95.6% ↗epoch run+2.7 ptClaude Opus 4.8
SAGE (Vals AI)54.8% ↗49.9% ↗+4.9 ptClaude Opus 4.8
SimpleBench (SimpleBench)64.8% ↗76.7% ↗-11.9 ptGemini 3.5 Flash
SimpleQA Verified53% ↗epoch run66.2% ↗epoch run-13.2 ptGemini 3.5 Flash
Terminal-Bench 4.0 (Vals AI)23.2% ↗6.1% ↗+17.1 ptClaude Opus 4.8
Toolathlon-Verified (HKUST)76.2% ↗67.3% ↗+8.9 ptClaude Opus 4.8
Vending-Bench 2 (Andon Labs)5787.43 ↗5396.42 ↗+391Claude Opus 4.8
WeirdML (Håvard Tveit Ihle)82.9% ↗62.6% ↗+20.3 ptClaude Opus 4.8
Only Claude Opus 4.8 reports · 12
SWE-bench Verified88.6% ↗not reported——
Humanity's Last Exam (no tools)49.8% ↗not reported——
BrowseComp84.3% ↗not reported——
Furniture Assembly (Epoch AI run)42.5% ↗epoch runnot reported——
GSO Opt@1 (GSO)47.1% ↗not reported——
Humanity's Last Exam (with tools)57.9% ↗not reported——
LMArena Agent (LMArena)0.0664 ↗not reported——
SWE-bench Multilingual84.4% ↗not reported——
SWE-bench Multimodal38.4% ↗not reported——
SWE-Bench Pro69.2% ↗not reported——
tau2-bench Banking Knowledge (Sierra)39.7% ↗not reported——
Terminal-Bench 2.174.6% ↗not reported——
Only Gemini 3.5 Flash reports · 8
ARC-AGI-2not reported72.1% ↗——
GDPVal-AAnot reported1656 ↗——
Humanity's Last Exam (full set, text + MM)not reported40.2% ↗——
MCP Atlasnot reported83.6% ↗——
MMMU-Pronot reported83.6% ↗——
SWE-Bench Pro (Public)not reported55.1% ↗——
SWE-bench Verified (Epoch AI run)not reported79.3% ↗epoch run——
Terminal-bench 2.1not reported76.2% ↗——
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Claude Opus 4.8 minus Gemini 3.5 Flash in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing

Side by side

SpecClaude Opus 4.8Gemini 3.5 FlashEdge
Context window1M tokens1.0M tokensGemini 3.5 Flash
Max output128K tokens66K tokensClaude Opus 4.8
Input price / 1M$5 ↗$1.5 ↗Gemini 3.5 Flash
Output price / 1M$25 ↗$9 ↗Gemini 3.5 Flash
Cached input / 1M$0.5 ↗$0.15 ↗Gemini 3.5 Flash
Input + output / 1M
Lower is cheaper. List prices; batch, tool and regional fees excluded.
$30$10.5Gemini 3.5 Flash
Modalitiestext · visiontext · vision · audioGemini 3.5 Flash
Released2026-05-282026-05-19—
Cited benchmark scores4742—
Reliability

Provider status

All providers →
More matchups

Claude Opus 4.8 vs …

More matchups

Gemini 3.5 Flash vs …

Built by Respan
Which one wins on your data?

Public benchmarks are a starting point. Run Claude Opus 4.8 and Gemini 3.5 Flash on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.