Compare / head-to-head

Claude Sonnet 5.5vsGemini 3.8 Flash

Claude Sonnet 5.5 leads 17 of 23 shared benchmarks. Gemini 3.8 Flash is 2.7x cheaper per token. Gemini 3.8 Flash has the larger context window (1.0M tokens).

Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.

Shared benchmarks
17 – 6
Claude Sonnet 5.5 leads
Cheaper per token
Gemini 3.8 Flash
2.7x cheaper, input + output
Larger context
Gemini 3.8 Flash
1.0M tokens
Providers
2 providers
Anthropic · Google
Claude Sonnet 5.5
Anthropic · released 2026-09-28
textvision
Context
1M
Max out
128K
Input /1M
$2 ↗
Output /1M
$10 ↗
Cached /1M
$0.2
Scores
29 · 10 core
Gemini 3.8 Flash
Google · released 2026-09-02
textvisionaudio
Context
1.0M
Max out
66K
Input /1M
$0.75 ↗
Output /1M
$3.75 ↗
Cached /1M
$0.075
Scores
43 · 10 core
Quality

Benchmark matrix

BenchmarkClaude Sonnet 5.5Gemini 3.8 FlashΔEdge
Reported by both · 23
GPQA Diamond95.6% ↗epoch run95.3% ↗+0.3 ptClaude Sonnet 5.5
LMArena Elo1471 ↗1494.8 ↗-23.8Gemini 3.8 Flash
APEX-Agents (Mercor)75.5% ↗64.3% ↗+11.2 ptClaude Sonnet 5.5
FrontierMath Tier 4 v2 (Epoch AI run)80.5% ↗epoch run22% ↗epoch run+58.5 ptClaude Sonnet 5.5
FrontierMath Tiers 1-3 v2 (Epoch AI run)88.8% ↗epoch run68.4% ↗epoch run+20.4 ptClaude Sonnet 5.5
FrontierSWE V2 (Proximal Labs)61.9% ↗19.6% ↗+42.3 ptClaude Sonnet 5.5
Furniture Assembly (Epoch AI run)75% ↗epoch run31.7% ↗epoch run+43.3 ptClaude Sonnet 5.5
LiveBench Agentic Coding (LiveBench)56.3% ↗54.2% ↗+2.1 ptClaude Sonnet 5.5
LiveBench Coding (LiveBench)91.4% ↗72.5% ↗+18.9 ptClaude Sonnet 5.5
LiveBench Data Analysis (LiveBench)78.6% ↗54% ↗+24.6 ptClaude Sonnet 5.5
LiveBench Instruction Following (LiveBench)70.5% ↗81.4% ↗-10.9 ptGemini 3.8 Flash
LiveBench Language (LiveBench)83.4% ↗87.8% ↗-4.4 ptGemini 3.8 Flash
LiveBench Mathematics (LiveBench)96.7% ↗91.6% ↗+5.1 ptClaude Sonnet 5.5
LiveBench Reasoning (LiveBench)91.6% ↗89.3% ↗+2.3 ptClaude Sonnet 5.5
LMArena Agent (LMArena)0.1252 ↗0.0297 ↗+0.1Claude Sonnet 5.5
LMArena Vision (LMArena)1268.3 ↗1290.3 ↗-22Gemini 3.8 Flash
LMArena WebDev (LMArena)1786.3 ↗1582.7 ↗+203.6Claude Sonnet 5.5
Mystery Game Puzzles (Epoch AI run)65% ↗epoch run47% ↗epoch run+18 ptClaude Sonnet 5.5
OTIS Mock AIME 2024-2025 (Epoch AI run)100% ↗epoch run98.9% ↗epoch run+1.1 ptClaude Sonnet 5.5
SAGE (Vals AI)51.8% ↗35.1% ↗+16.7 ptClaude Sonnet 5.5
SimpleBench (SimpleBench)75.9% ↗82.4% ↗-6.5 ptGemini 3.8 Flash
SimpleQA Verified46.5% ↗epoch run69.7% ↗epoch run-23.2 ptGemini 3.8 Flash
Terminal-Bench 4.0 (Vals AI)64.1% ↗19.2% ↗+44.9 ptClaude Sonnet 5.5
Only Claude Sonnet 5.5 reports · 6
Chartography (no tools)61.6% ↗not reported——
CursorBench 4.055.5% ↗not reported——
FrontierCode v1.1 (Main)46.2% ↗not reported——
Humanity's Last Exam (with tools)64.5% ↗not reported——
OSWorld 2.1 (partial)80.1% ↗not reported——
Terminal-Bench 4.070.6% ↗not reported——
Only Gemini 3.8 Flash reports · 14
CharXiv Reasoning (no tools)not reported86.2% ↗——
Chess Puzzles (Epoch AI run)not reported61% ↗epoch run——
DeepSWE v1.1not reported73.7% ↗——
DeepSWE v1.1 (Datacurve)not reported73.8% ↗——
GDPVal-AA v2not reported1545 ↗——
Harvey's Legal Agent Benchmarknot reported10% ↗——
HLE-Verifiednot reported54.9% ↗——
LABBench2not reported86.2% ↗——
OSWorld-2.0not reported59% ↗——
Terminal-bench 2.1not reported89.4% ↗——
Terminal-bench 4.0not reported19.1% ↗——
Vals Finance Agent v2not reported61.4% ↗——
Vending-Bench 2 (Andon Labs)not reported5093.79 ↗——
WeirdML (Håvard Tveit Ihle)not reported84.8% ↗——
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Claude Sonnet 5.5 minus Gemini 3.8 Flash in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing

Side by side

SpecClaude Sonnet 5.5Gemini 3.8 FlashEdge
Context window1M tokens1.0M tokensGemini 3.8 Flash
Max output128K tokens66K tokensClaude Sonnet 5.5
Input price / 1M$2 ↗$0.75 ↗Gemini 3.8 Flash
Output price / 1M$10 ↗$3.75 ↗Gemini 3.8 Flash
Cached input / 1M$0.2 ↗$0.075 ↗Gemini 3.8 Flash
Input + output / 1M
Lower is cheaper. List prices; batch, tool and regional fees excluded.
$12$4.5Gemini 3.8 Flash
Modalitiestext · visiontext · vision · audioGemini 3.8 Flash
Released2026-09-282026-09-02—
Cited benchmark scores2943—
Reliability

Provider status

All providers →
More matchups

Claude Sonnet 5.5 vs …

More matchups

Gemini 3.8 Flash vs …

Built by Respan
Which one wins on your data?

Public benchmarks are a starting point. Run Claude Sonnet 5.5 and Gemini 3.8 Flash on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.