Compare / head-to-head

Gemini 3.1 Pro PreviewvsGrok 4.5

Gemini 3.1 Pro Preview leads 16 of 26 shared benchmarks. Grok 4.5 is 1.8x cheaper per token. Gemini 3.1 Pro Preview has the larger context window (1.0M tokens).

Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.

Shared benchmarks
16 – 10
Gemini 3.1 Pro Preview leads
Cheaper per token
Grok 4.5
1.8x cheaper, input + output
Larger context
Gemini 3.1 Pro Preview
1.0M tokens
Providers
2 providers
Google · xAI
Gemini 3.1 Pro Preview
Google · released 2026-02-19
textvisionaudio
Context
1.0M
Max out
66K
Input /1M
$2 ↗
Output /1M
$12 ↗
Cached /1M
$0.2
Scores
43 · 12 core
Grok 4.5
xAI · released 2026-07-08
textvision
Context
500K
Max out
—
Input /1M
$2 ↗
Output /1M
$6 ↗
Cached /1M
$0.3
Scores
33 · 9 core
Quality

Benchmark matrix

BenchmarkGemini 3.1 Pro PreviewGrok 4.5ΔEdge
Reported by both · 26
GPQA Diamond94.3% ↗93.4% ↗epoch run+0.9 ptGemini 3.1 Pro Preview
LMArena Elo1480.1 ↗1450.1 ↗+30Gemini 3.1 Pro Preview
APEX-Agents (Mercor)35.3% ↗56.2% ↗-20.9 ptGrok 4.5
Chess Puzzles (Epoch AI run)55% ↗epoch run36% ↗epoch run+19 ptGemini 3.1 Pro Preview
DeepSWE v1.1 (Datacurve)11.7% ↗53.8% ↗-42.1 ptGrok 4.5
FrontierMath Tier 4 v2 (Epoch AI run)26.8% ↗epoch run24.4% ↗epoch run+2.4 ptGemini 3.1 Pro Preview
FrontierMath Tiers 1-3 v2 (Epoch AI run)59.6% ↗epoch run57.2% ↗epoch run+2.4 ptGemini 3.1 Pro Preview
Furniture Assembly (Epoch AI run)26.7% ↗epoch run22.5% ↗epoch run+4.2 ptGemini 3.1 Pro Preview
LiveBench Agentic Coding (LiveBench)44.1% ↗56.5% ↗-12.4 ptGrok 4.5
LiveBench Coding (LiveBench)76.5% ↗68.6% ↗+7.9 ptGemini 3.1 Pro Preview
LiveBench Data Analysis (LiveBench)78.5% ↗73% ↗+5.5 ptGemini 3.1 Pro Preview
LiveBench Instruction Following (LiveBench)79.1% ↗71.5% ↗+7.6 ptGemini 3.1 Pro Preview
LiveBench Language (LiveBench)85.4% ↗82.8% ↗+2.6 ptGemini 3.1 Pro Preview
LiveBench Mathematics (LiveBench)91% ↗90.8% ↗+0.2 ptGemini 3.1 Pro Preview
LiveBench Reasoning (LiveBench)84% ↗87.2% ↗-3.2 ptGrok 4.5
LMArena Agent (LMArena)-0.0771 ↗0.0121 ↗-0.1Grok 4.5
LMArena Vision (LMArena)1279.3 ↗1279 ↗+0.3Gemini 3.1 Pro Preview
LMArena WebDev (LMArena)1446.2 ↗1551.8 ↗-105.6Grok 4.5
OTIS Mock AIME 2024-2025 (Epoch AI run)95.6% ↗epoch run97.8% ↗epoch run-2.2 ptGrok 4.5
SAGE (Vals AI)48.7% ↗35% ↗+13.7 ptGemini 3.1 Pro Preview
SimpleBench (SimpleBench)79.6% ↗70% ↗+9.6 ptGemini 3.1 Pro Preview
SimpleQA Verified73.5% ↗epoch run48.3% ↗epoch run+25.2 ptGemini 3.1 Pro Preview
tau2-bench Banking Knowledge (Sierra)26% ↗47.9% ↗-21.9 ptGrok 4.5
Terminal-Bench 4.0 (Vals AI)2.5% ↗8.6% ↗-6.1 ptGrok 4.5
Vending-Bench 2 (Andon Labs)3774.25 ↗3887.43 ↗-113.2Grok 4.5
WeirdML (Håvard Tveit Ihle)72.1% ↗46.4% ↗+25.7 ptGemini 3.1 Pro Preview
Only Gemini 3.1 Pro Preview reports · 10
AIME 202698.3% ↗matharena ⚠not reported——
ARC-AGI-277.1% ↗not reported——
BALROG (BALROG)57% ↗not reported——
EBR-bench (Epoch AI run)14.3% ↗epoch runnot reported——
GSO Opt@1 (GSO)21.6% ↗not reported——
MirrorCode (Epoch AI run)8.9% ↗epoch runnot reported——
Mystery Game Puzzles (Epoch AI run)34% ↗epoch runnot reported——
SWE-Bench Pro54.2% ↗not reported——
SWE-bench Verified (Epoch AI run)75.6% ↗epoch runnot reported——
Toolathlon-Verified (HKUST)61.1% ↗not reported——
Only Grok 4.5 reports · 1
SWE-rebench 2026-05-15 to 2026-07-01 (Nebius)not reported63.8% ↗——
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Gemini 3.1 Pro Preview minus Grok 4.5 in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing

Side by side

SpecGemini 3.1 Pro PreviewGrok 4.5Edge
Context window1.0M tokens500K tokensGemini 3.1 Pro Preview
Max output66K tokens——
Input price / 1M$2 ↗$2 ↗Tie
Output price / 1M$12 ↗$6 ↗Grok 4.5
Cached input / 1M$0.2 ↗$0.3 ↗Gemini 3.1 Pro Preview
Input + output / 1M
Lower is cheaper. List prices; batch, tool and regional fees excluded.
$14$8Grok 4.5
Modalitiestext · vision · audiotext · visionGemini 3.1 Pro Preview
Released2026-02-192026-07-08—
Cited benchmark scores4333—
Reliability

Provider status

All providers →
More matchups

Gemini 3.1 Pro Preview vs …

Built by Respan
Which one wins on your data?

Public benchmarks are a starting point. Run Gemini 3.1 Pro Preview and Grok 4.5 on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.