Compare / head-to-head
Claude Opus 4.8vs
Gemini 3.5 Flash
Claude Opus 4.8 leads 22 of 28 shared benchmarks. Gemini 3.5 Flash is 2.9x cheaper per token. Gemini 3.5 Flash has the larger context window (1.0M tokens).
Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.
Shared benchmarks
22 – 6
Claude Opus 4.8 leads
Cheaper per token
Gemini 3.5 Flash
2.9x cheaper, input + output
Larger context
Gemini 3.5 Flash
1.0M tokens
Providers
2 providers
Anthropic · Google
Claude Opus 4.8
Anthropic · released 2026-05-28
textvision
Gemini 3.5 Flash
Google · released 2026-05-19
textvisionaudio
Quality
Benchmark matrix
28 shared · 12 only Claude Opus 4.8 · 8 only Gemini 3.5 Flash| Benchmark | Claude Opus 4.8 | Gemini 3.5 Flash | Δ | Edge |
|---|---|---|---|---|
| Reported by both · 28 | ||||
| GPQA Diamond | 93.6% ↗ | 92.8% ↗epoch run | +0.8 pt | Claude Opus 4.8 |
| AIME 2026 | 100% ↗matharena ⚠ | 95% ↗matharena ⚠ | +5 pt | Claude Opus 4.8 |
| LMArena Elo | 1452.7 ↗ | 1477.4 ↗ | -24.7 | Gemini 3.5 Flash |
| APEX-Agents (Mercor) | 48.9% ↗ | 27.5% ↗ | +21.4 pt | Claude Opus 4.8 |
| Chess Puzzles (Epoch AI run) | 34% ↗epoch run | 50% ↗epoch run | -16 pt | Gemini 3.5 Flash |
| DeepSWE v1.1 (Datacurve) | 59% ↗ | 36.1% ↗ | +22.9 pt | Claude Opus 4.8 |
| EBR-bench (Epoch AI run) | 28.6% ↗epoch run | 4.8% ↗epoch run | +23.8 pt | Claude Opus 4.8 |
| FrontierMath Tier 4 v2 (Epoch AI run) | 56.1% ↗epoch run | 26.8% ↗epoch run | +29.3 pt | Claude Opus 4.8 |
| FrontierMath Tiers 1-3 v2 (Epoch AI run) | 80% ↗epoch run | 62.8% ↗epoch run | +17.2 pt | Claude Opus 4.8 |
| LiveBench Agentic Coding (LiveBench) | 50.5% ↗ | 49% ↗ | +1.5 pt | Claude Opus 4.8 |
| LiveBench Coding (LiveBench) | 81.8% ↗ | 78.2% ↗ | +3.6 pt | Claude Opus 4.8 |
| LiveBench Data Analysis (LiveBench) | 66% ↗ | 64.9% ↗ | +1.1 pt | Claude Opus 4.8 |
| LiveBench Instruction Following (LiveBench) | 72% ↗ | 75.6% ↗ | -3.6 pt | Gemini 3.5 Flash |
| LiveBench Language (LiveBench) | 79.7% ↗ | 84.6% ↗ | -4.9 pt | Gemini 3.5 Flash |
| LiveBench Mathematics (LiveBench) | 94.3% ↗ | 88.2% ↗ | +6.1 pt | Claude Opus 4.8 |
| LiveBench Reasoning (LiveBench) | 89.2% ↗ | 82% ↗ | +7.2 pt | Claude Opus 4.8 |
| LMArena Vision (LMArena) | 1286.4 ↗ | 1284.4 ↗ | +2 | Claude Opus 4.8 |
| LMArena WebDev (LMArena) | 1555.5 ↗ | 1499 ↗ | +56.5 | Claude Opus 4.8 |
| Mystery Game Puzzles (Epoch AI run) | 36% ↗epoch run | 32% ↗epoch run | +4 pt | Claude Opus 4.8 |
| OSWorld-Verified | 83.4% ↗ | 78.4% ↗ | +5 pt | Claude Opus 4.8 |
| OTIS Mock AIME 2024-2025 (Epoch AI run) | 98.3% ↗epoch run | 95.6% ↗epoch run | +2.7 pt | Claude Opus 4.8 |
| SAGE (Vals AI) | 54.8% ↗ | 49.9% ↗ | +4.9 pt | Claude Opus 4.8 |
| SimpleBench (SimpleBench) | 64.8% ↗ | 76.7% ↗ | -11.9 pt | Gemini 3.5 Flash |
| SimpleQA Verified | 53% ↗epoch run | 66.2% ↗epoch run | -13.2 pt | Gemini 3.5 Flash |
| Terminal-Bench 4.0 (Vals AI) | 23.2% ↗ | 6.1% ↗ | +17.1 pt | Claude Opus 4.8 |
| Toolathlon-Verified (HKUST) | 76.2% ↗ | 67.3% ↗ | +8.9 pt | Claude Opus 4.8 |
| Vending-Bench 2 (Andon Labs) | 5787.43 ↗ | 5396.42 ↗ | +391 | Claude Opus 4.8 |
| WeirdML (Håvard Tveit Ihle) | 82.9% ↗ | 62.6% ↗ | +20.3 pt | Claude Opus 4.8 |
| Only Claude Opus 4.8 reports · 12 | ||||
| SWE-bench Verified | 88.6% ↗ | not reported | — | — |
| Humanity's Last Exam (no tools) | 49.8% ↗ | not reported | — | — |
| BrowseComp | 84.3% ↗ | not reported | — | — |
| Furniture Assembly (Epoch AI run) | 42.5% ↗epoch run | not reported | — | — |
| GSO Opt@1 (GSO) | 47.1% ↗ | not reported | — | — |
| Humanity's Last Exam (with tools) | 57.9% ↗ | not reported | — | — |
| LMArena Agent (LMArena) | 0.0664 ↗ | not reported | — | — |
| SWE-bench Multilingual | 84.4% ↗ | not reported | — | — |
| SWE-bench Multimodal | 38.4% ↗ | not reported | — | — |
| SWE-Bench Pro | 69.2% ↗ | not reported | — | — |
| tau2-bench Banking Knowledge (Sierra) | 39.7% ↗ | not reported | — | — |
| Terminal-Bench 2.1 | 74.6% ↗ | not reported | — | — |
| Only Gemini 3.5 Flash reports · 8 | ||||
| ARC-AGI-2 | not reported | 72.1% ↗ | — | — |
| GDPVal-AA | not reported | 1656 ↗ | — | — |
| Humanity's Last Exam (full set, text + MM) | not reported | 40.2% ↗ | — | — |
| MCP Atlas | not reported | 83.6% ↗ | — | — |
| MMMU-Pro | not reported | 83.6% ↗ | — | — |
| SWE-Bench Pro (Public) | not reported | 55.1% ↗ | — | — |
| SWE-bench Verified (Epoch AI run) | not reported | 79.3% ↗epoch run | — | — |
| Terminal-bench 2.1 | not reported | 76.2% ↗ | — | — |
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Claude Opus 4.8 minus Gemini 3.5 Flash in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing
Side by side
Official pricing pages and model cards| Spec | Claude Opus 4.8 | Gemini 3.5 Flash | Edge |
|---|---|---|---|
| Context window | 1M tokens | 1.0M tokens | Gemini 3.5 Flash |
| Max output | 128K tokens | 66K tokens | Claude Opus 4.8 |
| Input price / 1M | $5 ↗ | $1.5 ↗ | Gemini 3.5 Flash |
| Output price / 1M | $25 ↗ | $9 ↗ | Gemini 3.5 Flash |
| Cached input / 1M | $0.5 ↗ | $0.15 ↗ | Gemini 3.5 Flash |
| Input + output / 1M Lower is cheaper. List prices; batch, tool and regional fees excluded. | $30 | $10.5 | Gemini 3.5 Flash |
| Modalities | text · vision | text · vision · audio | Gemini 3.5 Flash |
| Released | 2026-05-28 | 2026-05-19 | — |
| Cited benchmark scores | 47 | 42 | — |
Reliability
Provider status
Live from /statusMore matchups
Claude Opus 4.8 vs …
Models sharing the most benchmarksMore matchups
Gemini 3.5 Flash vs …
Models sharing the most benchmarksBuilt by Respan
Which one wins on your data?
Public benchmarks are a starting point. Run Claude Opus 4.8 and Gemini 3.5 Flash on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.