Compare / head-to-head
Claude Sonnet 5.5vs
Gemini 3.8 Flash
Claude Sonnet 5.5 leads 17 of 23 shared benchmarks. Gemini 3.8 Flash is 2.7x cheaper per token. Gemini 3.8 Flash has the larger context window (1.0M tokens).
Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.
Shared benchmarks
17 – 6
Claude Sonnet 5.5 leads
Cheaper per token
Gemini 3.8 Flash
2.7x cheaper, input + output
Larger context
Gemini 3.8 Flash
1.0M tokens
Providers
2 providers
Anthropic · Google
Claude Sonnet 5.5
Anthropic · released 2026-09-28
textvision
Gemini 3.8 Flash
Google · released 2026-09-02
textvisionaudio
Quality
Benchmark matrix
23 shared · 6 only Claude Sonnet 5.5 · 14 only Gemini 3.8 Flash| Benchmark | Claude Sonnet 5.5 | Gemini 3.8 Flash | Δ | Edge |
|---|---|---|---|---|
| Reported by both · 23 | ||||
| GPQA Diamond | 95.6% ↗epoch run | 95.3% ↗ | +0.3 pt | Claude Sonnet 5.5 |
| LMArena Elo | 1471 ↗ | 1494.8 ↗ | -23.8 | Gemini 3.8 Flash |
| APEX-Agents (Mercor) | 75.5% ↗ | 64.3% ↗ | +11.2 pt | Claude Sonnet 5.5 |
| FrontierMath Tier 4 v2 (Epoch AI run) | 80.5% ↗epoch run | 22% ↗epoch run | +58.5 pt | Claude Sonnet 5.5 |
| FrontierMath Tiers 1-3 v2 (Epoch AI run) | 88.8% ↗epoch run | 68.4% ↗epoch run | +20.4 pt | Claude Sonnet 5.5 |
| FrontierSWE V2 (Proximal Labs) | 61.9% ↗ | 19.6% ↗ | +42.3 pt | Claude Sonnet 5.5 |
| Furniture Assembly (Epoch AI run) | 75% ↗epoch run | 31.7% ↗epoch run | +43.3 pt | Claude Sonnet 5.5 |
| LiveBench Agentic Coding (LiveBench) | 56.3% ↗ | 54.2% ↗ | +2.1 pt | Claude Sonnet 5.5 |
| LiveBench Coding (LiveBench) | 91.4% ↗ | 72.5% ↗ | +18.9 pt | Claude Sonnet 5.5 |
| LiveBench Data Analysis (LiveBench) | 78.6% ↗ | 54% ↗ | +24.6 pt | Claude Sonnet 5.5 |
| LiveBench Instruction Following (LiveBench) | 70.5% ↗ | 81.4% ↗ | -10.9 pt | Gemini 3.8 Flash |
| LiveBench Language (LiveBench) | 83.4% ↗ | 87.8% ↗ | -4.4 pt | Gemini 3.8 Flash |
| LiveBench Mathematics (LiveBench) | 96.7% ↗ | 91.6% ↗ | +5.1 pt | Claude Sonnet 5.5 |
| LiveBench Reasoning (LiveBench) | 91.6% ↗ | 89.3% ↗ | +2.3 pt | Claude Sonnet 5.5 |
| LMArena Agent (LMArena) | 0.1252 ↗ | 0.0297 ↗ | +0.1 | Claude Sonnet 5.5 |
| LMArena Vision (LMArena) | 1268.3 ↗ | 1290.3 ↗ | -22 | Gemini 3.8 Flash |
| LMArena WebDev (LMArena) | 1786.3 ↗ | 1582.7 ↗ | +203.6 | Claude Sonnet 5.5 |
| Mystery Game Puzzles (Epoch AI run) | 65% ↗epoch run | 47% ↗epoch run | +18 pt | Claude Sonnet 5.5 |
| OTIS Mock AIME 2024-2025 (Epoch AI run) | 100% ↗epoch run | 98.9% ↗epoch run | +1.1 pt | Claude Sonnet 5.5 |
| SAGE (Vals AI) | 51.8% ↗ | 35.1% ↗ | +16.7 pt | Claude Sonnet 5.5 |
| SimpleBench (SimpleBench) | 75.9% ↗ | 82.4% ↗ | -6.5 pt | Gemini 3.8 Flash |
| SimpleQA Verified | 46.5% ↗epoch run | 69.7% ↗epoch run | -23.2 pt | Gemini 3.8 Flash |
| Terminal-Bench 4.0 (Vals AI) | 64.1% ↗ | 19.2% ↗ | +44.9 pt | Claude Sonnet 5.5 |
| Only Claude Sonnet 5.5 reports · 6 | ||||
| Chartography (no tools) | 61.6% ↗ | not reported | — | — |
| CursorBench 4.0 | 55.5% ↗ | not reported | — | — |
| FrontierCode v1.1 (Main) | 46.2% ↗ | not reported | — | — |
| Humanity's Last Exam (with tools) | 64.5% ↗ | not reported | — | — |
| OSWorld 2.1 (partial) | 80.1% ↗ | not reported | — | — |
| Terminal-Bench 4.0 | 70.6% ↗ | not reported | — | — |
| Only Gemini 3.8 Flash reports · 14 | ||||
| CharXiv Reasoning (no tools) | not reported | 86.2% ↗ | — | — |
| Chess Puzzles (Epoch AI run) | not reported | 61% ↗epoch run | — | — |
| DeepSWE v1.1 | not reported | 73.7% ↗ | — | — |
| DeepSWE v1.1 (Datacurve) | not reported | 73.8% ↗ | — | — |
| GDPVal-AA v2 | not reported | 1545 ↗ | — | — |
| Harvey's Legal Agent Benchmark | not reported | 10% ↗ | — | — |
| HLE-Verified | not reported | 54.9% ↗ | — | — |
| LABBench2 | not reported | 86.2% ↗ | — | — |
| OSWorld-2.0 | not reported | 59% ↗ | — | — |
| Terminal-bench 2.1 | not reported | 89.4% ↗ | — | — |
| Terminal-bench 4.0 | not reported | 19.1% ↗ | — | — |
| Vals Finance Agent v2 | not reported | 61.4% ↗ | — | — |
| Vending-Bench 2 (Andon Labs) | not reported | 5093.79 ↗ | — | — |
| WeirdML (Håvard Tveit Ihle) | not reported | 84.8% ↗ | — | — |
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Claude Sonnet 5.5 minus Gemini 3.8 Flash in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing
Side by side
Official pricing pages and model cards| Spec | Claude Sonnet 5.5 | Gemini 3.8 Flash | Edge |
|---|---|---|---|
| Context window | 1M tokens | 1.0M tokens | Gemini 3.8 Flash |
| Max output | 128K tokens | 66K tokens | Claude Sonnet 5.5 |
| Input price / 1M | $2 ↗ | $0.75 ↗ | Gemini 3.8 Flash |
| Output price / 1M | $10 ↗ | $3.75 ↗ | Gemini 3.8 Flash |
| Cached input / 1M | $0.2 ↗ | $0.075 ↗ | Gemini 3.8 Flash |
| Input + output / 1M Lower is cheaper. List prices; batch, tool and regional fees excluded. | $12 | $4.5 | Gemini 3.8 Flash |
| Modalities | text · vision | text · vision · audio | Gemini 3.8 Flash |
| Released | 2026-09-28 | 2026-09-02 | — |
| Cited benchmark scores | 29 | 43 | — |
Reliability
Provider status
Live from /statusMore matchups
Claude Sonnet 5.5 vs …
Models sharing the most benchmarksMore matchups
Gemini 3.8 Flash vs …
Models sharing the most benchmarksBuilt by Respan
Which one wins on your data?
Public benchmarks are a starting point. Run Claude Sonnet 5.5 and Gemini 3.8 Flash on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.