Compare / head-to-head

DeepSeek V4 FlashvsGemini 3.1 Pro Preview

Gemini 3.1 Pro Preview leads 13 of 21 shared benchmarks. Gemini 3.1 Pro Preview has the larger context window (1.0M tokens).

Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.

Shared benchmarks
7 – 13
Gemini 3.1 Pro Preview leads · 1 tied
Cheaper per token
—
list price, input + output
Larger context
Gemini 3.1 Pro Preview
1.0M tokens
Providers
2 providers
DeepSeek · Google
DeepSeek V4 Flash
DeepSeek · released 2026-07-31
text
Context
1M
Max out
—
Input /1M
—
Output /1M
—
Cached /1M
—
Scores
28 · 7 core
Gemini 3.1 Pro Preview
Google · released 2026-02-19
textvisionaudio
Context
1.0M
Max out
66K
Input /1M
$2 ↗
Output /1M
$12 ↗
Cached /1M
$0.2
Scores
43 · 12 core
Quality

Benchmark matrix

BenchmarkDeepSeek V4 FlashGemini 3.1 Pro PreviewΔEdge
Reported by both · 21
GPQA Diamond91% ↗epoch run94.3% ↗-3.3 ptGemini 3.1 Pro Preview
LMArena Elo1436 ↗1480.1 ↗-44.1Gemini 3.1 Pro Preview
Chess Puzzles (Epoch AI run)33% ↗epoch run55% ↗epoch run-22 ptGemini 3.1 Pro Preview
DeepSWE v1.1 (Datacurve)53.3% ↗11.7% ↗+41.6 ptDeepSeek V4 Flash
FrontierMath Tier 4 v2 (Epoch AI run)24.4% ↗epoch run26.8% ↗epoch run-2.4 ptGemini 3.1 Pro Preview
FrontierMath Tiers 1-3 v2 (Epoch AI run)57.5% ↗epoch run59.6% ↗epoch run-2.1 ptGemini 3.1 Pro Preview
LiveBench Agentic Coding (LiveBench)46.8% ↗44.1% ↗+2.7 ptDeepSeek V4 Flash
LiveBench Coding (LiveBench)75% ↗76.5% ↗-1.5 ptGemini 3.1 Pro Preview
LiveBench Data Analysis (LiveBench)79.3% ↗78.5% ↗+0.8 ptDeepSeek V4 Flash
LiveBench Instruction Following (LiveBench)65.5% ↗79.1% ↗-13.6 ptGemini 3.1 Pro Preview
LiveBench Language (LiveBench)79.2% ↗85.4% ↗-6.2 ptGemini 3.1 Pro Preview
LiveBench Mathematics (LiveBench)86.8% ↗91% ↗-4.2 ptGemini 3.1 Pro Preview
LiveBench Reasoning (LiveBench)86.6% ↗84% ↗+2.6 ptDeepSeek V4 Flash
LMArena WebDev (LMArena)1581 ↗1446.2 ↗+134.8DeepSeek V4 Flash
Mystery Game Puzzles (Epoch AI run)34% ↗epoch run34% ↗epoch run0Tie
OTIS Mock AIME 2024-2025 (Epoch AI run)94.4% ↗epoch run95.6% ↗epoch run-1.2 ptGemini 3.1 Pro Preview
SimpleBench (SimpleBench)61.1% ↗79.6% ↗-18.5 ptGemini 3.1 Pro Preview
SimpleQA Verified33.6% ↗epoch run73.5% ↗epoch run-39.9 ptGemini 3.1 Pro Preview
Terminal-Bench 4.0 (Vals AI)18.7% ↗2.5% ↗+16.2 ptDeepSeek V4 Flash
Toolathlon-Verified (HKUST)70.7% ↗61.1% ↗+9.6 ptDeepSeek V4 Flash
WeirdML (Håvard Tveit Ihle)63% ↗72.1% ↗-9.1 ptGemini 3.1 Pro Preview
Only DeepSeek V4 Flash reports · 7
Agents' Last Exam25.2% ↗not reported——
AutomationBench Public25.1% ↗not reported——
Cybergym76.7% ↗not reported——
DeepSWE54.4% ↗not reported——
NL2Repo54.2% ↗not reported——
Terminal Bench 2.182.7% ↗not reported——
Toolathlon-Verified70.3% ↗not reported——
Only Gemini 3.1 Pro Preview reports · 15
AIME 2026not reported98.3% ↗matharena ⚠——
APEX-Agents (Mercor)not reported35.3% ↗——
ARC-AGI-2not reported77.1% ↗——
BALROG (BALROG)not reported57% ↗——
EBR-bench (Epoch AI run)not reported14.3% ↗epoch run——
Furniture Assembly (Epoch AI run)not reported26.7% ↗epoch run——
GSO Opt@1 (GSO)not reported21.6% ↗——
LMArena Agent (LMArena)not reported-0.0771 ↗——
LMArena Vision (LMArena)not reported1279.3 ↗——
MirrorCode (Epoch AI run)not reported8.9% ↗epoch run——
SAGE (Vals AI)not reported48.7% ↗——
SWE-Bench Pronot reported54.2% ↗——
SWE-bench Verified (Epoch AI run)not reported75.6% ↗epoch run——
tau2-bench Banking Knowledge (Sierra)not reported26% ↗——
Vending-Bench 2 (Andon Labs)not reported3774.25 ↗——
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is DeepSeek V4 Flash minus Gemini 3.1 Pro Preview in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing

Side by side

SpecDeepSeek V4 FlashGemini 3.1 Pro PreviewEdge
Context window1M tokens1.0M tokensGemini 3.1 Pro Preview
Max output—66K tokens—
Input price / 1M—$2 ↗—
Output price / 1M—$12 ↗—
Cached input / 1M—$0.2 ↗—
Input + output / 1M
Lower is cheaper. List prices; batch, tool and regional fees excluded.
—$14—
Modalitiestexttext · vision · audioGemini 3.1 Pro Preview
Released2026-07-312026-02-19—
Cited benchmark scores2843—
Reliability

Provider status

All providers →
More matchups

DeepSeek V4 Flash vs …

More matchups

Gemini 3.1 Pro Preview vs …

Built by Respan
Which one wins on your data?

Public benchmarks are a starting point. Run DeepSeek V4 Flash and Gemini 3.1 Pro Preview on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.