Compare / head-to-head
Gemini 3.8 Flashvs
Grok 4.5
Gemini 3.8 Flash leads 21 of 25 shared benchmarks. Gemini 3.8 Flash is 1.8x cheaper per token. Gemini 3.8 Flash has the larger context window (1.0M tokens).
Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.
Shared benchmarks
21 – 3
Gemini 3.8 Flash leads · 1 tied
Cheaper per token
Gemini 3.8 Flash
1.8x cheaper, input + output
Larger context
Gemini 3.8 Flash
1.0M tokens
Providers
2 providers
Google · xAI
Gemini 3.8 Flash
Google · released 2026-09-02
textvisionaudio
Quality
Benchmark matrix
25 shared · 12 only Gemini 3.8 Flash · 2 only Grok 4.5| Benchmark | Gemini 3.8 Flash | Grok 4.5 | Δ | Edge |
|---|---|---|---|---|
| Reported by both · 25 | ||||
| GPQA Diamond | 95.3% ↗ | 93.4% ↗epoch run | +1.9 pt | Gemini 3.8 Flash |
| LMArena Elo | 1494.8 ↗ | 1450.1 ↗ | +44.7 | Gemini 3.8 Flash |
| APEX-Agents (Mercor) | 64.3% ↗ | 56.2% ↗ | +8.1 pt | Gemini 3.8 Flash |
| Chess Puzzles (Epoch AI run) | 61% ↗epoch run | 36% ↗epoch run | +25 pt | Gemini 3.8 Flash |
| DeepSWE v1.1 (Datacurve) | 73.8% ↗ | 53.8% ↗ | +20 pt | Gemini 3.8 Flash |
| FrontierMath Tier 4 v2 (Epoch AI run) | 22% ↗epoch run | 24.4% ↗epoch run | -2.4 pt | Grok 4.5 |
| FrontierMath Tiers 1-3 v2 (Epoch AI run) | 68.4% ↗epoch run | 57.2% ↗epoch run | +11.2 pt | Gemini 3.8 Flash |
| Furniture Assembly (Epoch AI run) | 31.7% ↗epoch run | 22.5% ↗epoch run | +9.2 pt | Gemini 3.8 Flash |
| LiveBench Agentic Coding (LiveBench) | 54.2% ↗ | 56.5% ↗ | -2.3 pt | Grok 4.5 |
| LiveBench Coding (LiveBench) | 72.5% ↗ | 68.6% ↗ | +3.9 pt | Gemini 3.8 Flash |
| LiveBench Data Analysis (LiveBench) | 54% ↗ | 73% ↗ | -19 pt | Grok 4.5 |
| LiveBench Instruction Following (LiveBench) | 81.4% ↗ | 71.5% ↗ | +9.9 pt | Gemini 3.8 Flash |
| LiveBench Language (LiveBench) | 87.8% ↗ | 82.8% ↗ | +5 pt | Gemini 3.8 Flash |
| LiveBench Mathematics (LiveBench) | 91.6% ↗ | 90.8% ↗ | +0.8 pt | Gemini 3.8 Flash |
| LiveBench Reasoning (LiveBench) | 89.3% ↗ | 87.2% ↗ | +2.1 pt | Gemini 3.8 Flash |
| LMArena Agent (LMArena) | 0.0297 ↗ | 0.0121 ↗ | 0 | Tie |
| LMArena Vision (LMArena) | 1290.3 ↗ | 1279 ↗ | +11.3 | Gemini 3.8 Flash |
| LMArena WebDev (LMArena) | 1582.7 ↗ | 1551.8 ↗ | +30.9 | Gemini 3.8 Flash |
| OTIS Mock AIME 2024-2025 (Epoch AI run) | 98.9% ↗epoch run | 97.8% ↗epoch run | +1.1 pt | Gemini 3.8 Flash |
| SAGE (Vals AI) | 35.1% ↗ | 35% ↗ | +0.1 pt | Gemini 3.8 Flash |
| SimpleBench (SimpleBench) | 82.4% ↗ | 70% ↗ | +12.4 pt | Gemini 3.8 Flash |
| SimpleQA Verified | 69.7% ↗epoch run | 48.3% ↗epoch run | +21.4 pt | Gemini 3.8 Flash |
| Terminal-Bench 4.0 (Vals AI) | 19.2% ↗ | 8.6% ↗ | +10.6 pt | Gemini 3.8 Flash |
| Vending-Bench 2 (Andon Labs) | 5093.79 ↗ | 3887.43 ↗ | +1206.4 | Gemini 3.8 Flash |
| WeirdML (Håvard Tveit Ihle) | 84.8% ↗ | 46.4% ↗ | +38.4 pt | Gemini 3.8 Flash |
| Only Gemini 3.8 Flash reports · 12 | ||||
| CharXiv Reasoning (no tools) | 86.2% ↗ | not reported | — | — |
| DeepSWE v1.1 | 73.7% ↗ | not reported | — | — |
| FrontierSWE V2 (Proximal Labs) | 19.6% ↗ | not reported | — | — |
| GDPVal-AA v2 | 1545 ↗ | not reported | — | — |
| Harvey's Legal Agent Benchmark | 10% ↗ | not reported | — | — |
| HLE-Verified | 54.9% ↗ | not reported | — | — |
| LABBench2 | 86.2% ↗ | not reported | — | — |
| Mystery Game Puzzles (Epoch AI run) | 47% ↗epoch run | not reported | — | — |
| OSWorld-2.0 | 59% ↗ | not reported | — | — |
| Terminal-bench 2.1 | 89.4% ↗ | not reported | — | — |
| Terminal-bench 4.0 | 19.1% ↗ | not reported | — | — |
| Vals Finance Agent v2 | 61.4% ↗ | not reported | — | — |
| Only Grok 4.5 reports · 2 | ||||
| SWE-rebench 2026-05-15 to 2026-07-01 (Nebius) | not reported | 63.8% ↗ | — | — |
| tau2-bench Banking Knowledge (Sierra) | not reported | 47.9% ↗ | — | — |
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Gemini 3.8 Flash minus Grok 4.5 in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing
Side by side
Official pricing pages and model cards| Spec | Gemini 3.8 Flash | Grok 4.5 | Edge |
|---|---|---|---|
| Context window | 1.0M tokens | 500K tokens | Gemini 3.8 Flash |
| Max output | 66K tokens | — | — |
| Input price / 1M | $0.75 ↗ | $2 ↗ | Gemini 3.8 Flash |
| Output price / 1M | $3.75 ↗ | $6 ↗ | Gemini 3.8 Flash |
| Cached input / 1M | $0.075 ↗ | $0.3 ↗ | Gemini 3.8 Flash |
| Input + output / 1M Lower is cheaper. List prices; batch, tool and regional fees excluded. | $4.5 | $8 | Gemini 3.8 Flash |
| Modalities | text · vision · audio | text · vision | Gemini 3.8 Flash |
| Released | 2026-09-02 | 2026-07-08 | — |
| Cited benchmark scores | 43 | 33 | — |
Reliability
Provider status
Live from /statusMore matchups
Gemini 3.8 Flash vs …
Models sharing the most benchmarksMore matchups
Grok 4.5 vs …
Models sharing the most benchmarksBuilt by Respan
Which one wins on your data?
Public benchmarks are a starting point. Run Gemini 3.8 Flash and Grok 4.5 on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.