Compare / head-to-head
DeepSeek V4 Provs
Gemini 3.1 Pro Preview
DeepSeek V4 Pro leads 14 of 26 shared benchmarks. DeepSeek V4 Pro is 2.7x cheaper per token. Gemini 3.1 Pro Preview has the larger context window (1.0M tokens).
Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.
Shared benchmarks
14 – 12
DeepSeek V4 Pro leads
Cheaper per token
DeepSeek V4 Pro
2.7x cheaper, input + output
Larger context
Gemini 3.1 Pro Preview
1.0M tokens
Providers
2 providers
DeepSeek · Google
DeepSeek V4 Pro
DeepSeek · released 2026-08-13
text
Gemini 3.1 Pro Preview
Google · released 2026-02-19
textvisionaudio
Quality
Benchmark matrix
26 shared · 4 only DeepSeek V4 Pro · 10 only Gemini 3.1 Pro Preview| Benchmark | DeepSeek V4 Pro | Gemini 3.1 Pro Preview | Δ | Edge |
|---|---|---|---|---|
| Reported by both · 26 | ||||
| GPQA Diamond | 90.9% ↗epoch run | 94.3% ↗ | -3.4 pt | Gemini 3.1 Pro Preview |
| AIME 2026 | 96.7% ↗matharena ⚠ | 98.3% ↗matharena ⚠ | -1.6 pt | Gemini 3.1 Pro Preview |
| LMArena Elo | 1450.6 ↗ | 1480.1 ↗ | -29.5 | Gemini 3.1 Pro Preview |
| APEX-Agents (Mercor) | 47.3% ↗ | 35.3% ↗ | +12 pt | DeepSeek V4 Pro |
| Chess Puzzles (Epoch AI run) | 47% ↗epoch run | 55% ↗epoch run | -8 pt | Gemini 3.1 Pro Preview |
| DeepSWE v1.1 (Datacurve) | 62.8% ↗ | 11.7% ↗ | +51.1 pt | DeepSeek V4 Pro |
| FrontierMath Tier 4 v2 (Epoch AI run) | 2.4% ↗epoch run | 26.8% ↗epoch run | -24.4 pt | Gemini 3.1 Pro Preview |
| FrontierMath Tiers 1-3 v2 (Epoch AI run) | 45.3% ↗epoch run | 59.6% ↗epoch run | -14.3 pt | Gemini 3.1 Pro Preview |
| LiveBench Agentic Coding (LiveBench) | 54.9% ↗ | 44.1% ↗ | +10.8 pt | DeepSeek V4 Pro |
| LiveBench Coding (LiveBench) | 77.2% ↗ | 76.5% ↗ | +0.7 pt | DeepSeek V4 Pro |
| LiveBench Data Analysis (LiveBench) | 79.2% ↗ | 78.5% ↗ | +0.7 pt | DeepSeek V4 Pro |
| LiveBench Instruction Following (LiveBench) | 67.7% ↗ | 79.1% ↗ | -11.4 pt | Gemini 3.1 Pro Preview |
| LiveBench Language (LiveBench) | 82.1% ↗ | 85.4% ↗ | -3.3 pt | Gemini 3.1 Pro Preview |
| LiveBench Mathematics (LiveBench) | 95.1% ↗ | 91% ↗ | +4.1 pt | DeepSeek V4 Pro |
| LiveBench Reasoning (LiveBench) | 85.8% ↗ | 84% ↗ | +1.8 pt | DeepSeek V4 Pro |
| LMArena Agent (LMArena) | 0.0119 ↗ | -0.0771 ↗ | +0.1 | DeepSeek V4 Pro |
| LMArena WebDev (LMArena) | 1582.6 ↗ | 1446.2 ↗ | +136.4 | DeepSeek V4 Pro |
| Mystery Game Puzzles (Epoch AI run) | 43% ↗epoch run | 34% ↗epoch run | +9 pt | DeepSeek V4 Pro |
| OTIS Mock AIME 2024-2025 (Epoch AI run) | 96.7% ↗epoch run | 95.6% ↗epoch run | +1.1 pt | DeepSeek V4 Pro |
| SimpleBench (SimpleBench) | 50.9% ↗ | 79.6% ↗ | -28.7 pt | Gemini 3.1 Pro Preview |
| SimpleQA Verified | 47% ↗epoch run | 73.5% ↗epoch run | -26.5 pt | Gemini 3.1 Pro Preview |
| SWE-bench Verified (Epoch AI run) | 77.6% ↗epoch run | 75.6% ↗epoch run | +2 pt | DeepSeek V4 Pro |
| Terminal-Bench 4.0 (Vals AI) | 14.1% ↗ | 2.5% ↗ | +11.6 pt | DeepSeek V4 Pro |
| Toolathlon-Verified (HKUST) | 74.4% ↗ | 61.1% ↗ | +13.3 pt | DeepSeek V4 Pro |
| Vending-Bench 2 (Andon Labs) | 3284.52 ↗ | 3774.25 ↗ | -489.7 | Gemini 3.1 Pro Preview |
| WeirdML (Håvard Tveit Ihle) | 66.2% ↗ | 72.1% ↗ | -5.9 pt | Gemini 3.1 Pro Preview |
| Only DeepSeek V4 Pro reports · 4 | ||||
| DeepSWE v1.1 | 62.7% ↗ | not reported | — | — |
| Humanity's Last Exam (with tools) | 60% ↗ | not reported | — | — |
| SWE-rebench 2026-05-15 to 2026-07-01 (Nebius) | 40.2% ↗ | not reported | — | — |
| Terminal-Bench 2.1 | 87.9% ↗ | not reported | — | — |
| Only Gemini 3.1 Pro Preview reports · 10 | ||||
| ARC-AGI-2 | not reported | 77.1% ↗ | — | — |
| BALROG (BALROG) | not reported | 57% ↗ | — | — |
| EBR-bench (Epoch AI run) | not reported | 14.3% ↗epoch run | — | — |
| Furniture Assembly (Epoch AI run) | not reported | 26.7% ↗epoch run | — | — |
| GSO Opt@1 (GSO) | not reported | 21.6% ↗ | — | — |
| LMArena Vision (LMArena) | not reported | 1279.3 ↗ | — | — |
| MirrorCode (Epoch AI run) | not reported | 8.9% ↗epoch run | — | — |
| SAGE (Vals AI) | not reported | 48.7% ↗ | — | — |
| SWE-Bench Pro | not reported | 54.2% ↗ | — | — |
| tau2-bench Banking Knowledge (Sierra) | not reported | 26% ↗ | — | — |
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is DeepSeek V4 Pro minus Gemini 3.1 Pro Preview in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing
Side by side
Official pricing pages and model cards| Spec | DeepSeek V4 Pro | Gemini 3.1 Pro Preview | Edge |
|---|---|---|---|
| Context window | 1M tokens | 1.0M tokens | Gemini 3.1 Pro Preview |
| Max output | 384K tokens | 66K tokens | DeepSeek V4 Pro |
| Input price / 1M | $1.32 ↗ | $2 ↗ | DeepSeek V4 Pro |
| Output price / 1M | $3.96 ↗ | $12 ↗ | DeepSeek V4 Pro |
| Cached input / 1M | $0.044 ↗ | $0.2 ↗ | DeepSeek V4 Pro |
| Input + output / 1M Lower is cheaper. List prices; batch, tool and regional fees excluded. | $5.28 | $14 | DeepSeek V4 Pro |
| Modalities | text | text · vision · audio | Gemini 3.1 Pro Preview |
| Released | 2026-08-13 | 2026-02-19 | — |
| Cited benchmark scores | 37 | 43 | — |
Reliability
Provider status
Live from /statusMore matchups
DeepSeek V4 Pro vs …
Models sharing the most benchmarksMore matchups
Gemini 3.1 Pro Preview vs …
Models sharing the most benchmarksBuilt by Respan
Which one wins on your data?
Public benchmarks are a starting point. Run DeepSeek V4 Pro and Gemini 3.1 Pro Preview on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.