Compare / head-to-head

Gemini 3.8 FlashvsGrok 4.5

Gemini 3.8 Flash leads 21 of 25 shared benchmarks. Gemini 3.8 Flash is 1.8x cheaper per token. Gemini 3.8 Flash has the larger context window (1.0M tokens).

Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.

Shared benchmarks
21 – 3
Gemini 3.8 Flash leads · 1 tied
Cheaper per token
Gemini 3.8 Flash
1.8x cheaper, input + output
Larger context
Gemini 3.8 Flash
1.0M tokens
Providers
2 providers
Google · xAI
Gemini 3.8 Flash
Google · released 2026-09-02
textvisionaudio
Context
1.0M
Max out
66K
Input /1M
$0.75 ↗
Output /1M
$3.75 ↗
Cached /1M
$0.075
Scores
43 · 10 core
Grok 4.5
xAI · released 2026-07-08
textvision
Context
500K
Max out
—
Input /1M
$2 ↗
Output /1M
$6 ↗
Cached /1M
$0.3
Scores
33 · 9 core
Quality

Benchmark matrix

BenchmarkGemini 3.8 FlashGrok 4.5ΔEdge
Reported by both · 25
GPQA Diamond95.3% ↗93.4% ↗epoch run+1.9 ptGemini 3.8 Flash
LMArena Elo1494.8 ↗1450.1 ↗+44.7Gemini 3.8 Flash
APEX-Agents (Mercor)64.3% ↗56.2% ↗+8.1 ptGemini 3.8 Flash
Chess Puzzles (Epoch AI run)61% ↗epoch run36% ↗epoch run+25 ptGemini 3.8 Flash
DeepSWE v1.1 (Datacurve)73.8% ↗53.8% ↗+20 ptGemini 3.8 Flash
FrontierMath Tier 4 v2 (Epoch AI run)22% ↗epoch run24.4% ↗epoch run-2.4 ptGrok 4.5
FrontierMath Tiers 1-3 v2 (Epoch AI run)68.4% ↗epoch run57.2% ↗epoch run+11.2 ptGemini 3.8 Flash
Furniture Assembly (Epoch AI run)31.7% ↗epoch run22.5% ↗epoch run+9.2 ptGemini 3.8 Flash
LiveBench Agentic Coding (LiveBench)54.2% ↗56.5% ↗-2.3 ptGrok 4.5
LiveBench Coding (LiveBench)72.5% ↗68.6% ↗+3.9 ptGemini 3.8 Flash
LiveBench Data Analysis (LiveBench)54% ↗73% ↗-19 ptGrok 4.5
LiveBench Instruction Following (LiveBench)81.4% ↗71.5% ↗+9.9 ptGemini 3.8 Flash
LiveBench Language (LiveBench)87.8% ↗82.8% ↗+5 ptGemini 3.8 Flash
LiveBench Mathematics (LiveBench)91.6% ↗90.8% ↗+0.8 ptGemini 3.8 Flash
LiveBench Reasoning (LiveBench)89.3% ↗87.2% ↗+2.1 ptGemini 3.8 Flash
LMArena Agent (LMArena)0.0297 ↗0.0121 ↗0Tie
LMArena Vision (LMArena)1290.3 ↗1279 ↗+11.3Gemini 3.8 Flash
LMArena WebDev (LMArena)1582.7 ↗1551.8 ↗+30.9Gemini 3.8 Flash
OTIS Mock AIME 2024-2025 (Epoch AI run)98.9% ↗epoch run97.8% ↗epoch run+1.1 ptGemini 3.8 Flash
SAGE (Vals AI)35.1% ↗35% ↗+0.1 ptGemini 3.8 Flash
SimpleBench (SimpleBench)82.4% ↗70% ↗+12.4 ptGemini 3.8 Flash
SimpleQA Verified69.7% ↗epoch run48.3% ↗epoch run+21.4 ptGemini 3.8 Flash
Terminal-Bench 4.0 (Vals AI)19.2% ↗8.6% ↗+10.6 ptGemini 3.8 Flash
Vending-Bench 2 (Andon Labs)5093.79 ↗3887.43 ↗+1206.4Gemini 3.8 Flash
WeirdML (Håvard Tveit Ihle)84.8% ↗46.4% ↗+38.4 ptGemini 3.8 Flash
Only Gemini 3.8 Flash reports · 12
CharXiv Reasoning (no tools)86.2% ↗not reported——
DeepSWE v1.173.7% ↗not reported——
FrontierSWE V2 (Proximal Labs)19.6% ↗not reported——
GDPVal-AA v21545 ↗not reported——
Harvey's Legal Agent Benchmark10% ↗not reported——
HLE-Verified54.9% ↗not reported——
LABBench286.2% ↗not reported——
Mystery Game Puzzles (Epoch AI run)47% ↗epoch runnot reported——
OSWorld-2.059% ↗not reported——
Terminal-bench 2.189.4% ↗not reported——
Terminal-bench 4.019.1% ↗not reported——
Vals Finance Agent v261.4% ↗not reported——
Only Grok 4.5 reports · 2
SWE-rebench 2026-05-15 to 2026-07-01 (Nebius)not reported63.8% ↗——
tau2-bench Banking Knowledge (Sierra)not reported47.9% ↗——
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Gemini 3.8 Flash minus Grok 4.5 in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing

Side by side

SpecGemini 3.8 FlashGrok 4.5Edge
Context window1.0M tokens500K tokensGemini 3.8 Flash
Max output66K tokens——
Input price / 1M$0.75 ↗$2 ↗Gemini 3.8 Flash
Output price / 1M$3.75 ↗$6 ↗Gemini 3.8 Flash
Cached input / 1M$0.075 ↗$0.3 ↗Gemini 3.8 Flash
Input + output / 1M
Lower is cheaper. List prices; batch, tool and regional fees excluded.
$4.5$8Gemini 3.8 Flash
Modalitiestext · vision · audiotext · visionGemini 3.8 Flash
Released2026-09-022026-07-08—
Cited benchmark scores4333—
Reliability

Provider status

All providers →
More matchups

Gemini 3.8 Flash vs …

Built by Respan
Which one wins on your data?

Public benchmarks are a starting point. Run Gemini 3.8 Flash and Grok 4.5 on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.