Compare / head-to-head
DeepSeek V4 Flashvs
Gemini 3.5 Flash
Gemini 3.5 Flash leads 13 of 21 shared benchmarks. Gemini 3.5 Flash has the larger context window (1.0M tokens).
Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.
Shared benchmarks
8 – 13
Gemini 3.5 Flash leads
Cheaper per token
—
list price, input + output
Larger context
Gemini 3.5 Flash
1.0M tokens
Providers
2 providers
DeepSeek · Google
DeepSeek V4 Flash
DeepSeek · released 2026-07-31
text
- Context
- 1M
- Max out
- —
- Input /1M
- —
- Output /1M
- —
- Cached /1M
- —
- Scores
- 28 · 7 core
Gemini 3.5 Flash
Google · released 2026-05-19
textvisionaudio
Quality
Benchmark matrix
21 shared · 7 only DeepSeek V4 Flash · 15 only Gemini 3.5 Flash| Benchmark | DeepSeek V4 Flash | Gemini 3.5 Flash | Δ | Edge |
|---|---|---|---|---|
| Reported by both · 21 | ||||
| GPQA Diamond | 91% ↗epoch run | 92.8% ↗epoch run | -1.8 pt | Gemini 3.5 Flash |
| LMArena Elo | 1436 ↗ | 1477.4 ↗ | -41.4 | Gemini 3.5 Flash |
| Chess Puzzles (Epoch AI run) | 33% ↗epoch run | 50% ↗epoch run | -17 pt | Gemini 3.5 Flash |
| DeepSWE v1.1 (Datacurve) | 53.3% ↗ | 36.1% ↗ | +17.2 pt | DeepSeek V4 Flash |
| FrontierMath Tier 4 v2 (Epoch AI run) | 24.4% ↗epoch run | 26.8% ↗epoch run | -2.4 pt | Gemini 3.5 Flash |
| FrontierMath Tiers 1-3 v2 (Epoch AI run) | 57.5% ↗epoch run | 62.8% ↗epoch run | -5.3 pt | Gemini 3.5 Flash |
| LiveBench Agentic Coding (LiveBench) | 46.8% ↗ | 49% ↗ | -2.2 pt | Gemini 3.5 Flash |
| LiveBench Coding (LiveBench) | 75% ↗ | 78.2% ↗ | -3.2 pt | Gemini 3.5 Flash |
| LiveBench Data Analysis (LiveBench) | 79.3% ↗ | 64.9% ↗ | +14.4 pt | DeepSeek V4 Flash |
| LiveBench Instruction Following (LiveBench) | 65.5% ↗ | 75.6% ↗ | -10.1 pt | Gemini 3.5 Flash |
| LiveBench Language (LiveBench) | 79.2% ↗ | 84.6% ↗ | -5.4 pt | Gemini 3.5 Flash |
| LiveBench Mathematics (LiveBench) | 86.8% ↗ | 88.2% ↗ | -1.4 pt | Gemini 3.5 Flash |
| LiveBench Reasoning (LiveBench) | 86.6% ↗ | 82% ↗ | +4.6 pt | DeepSeek V4 Flash |
| LMArena WebDev (LMArena) | 1581 ↗ | 1499 ↗ | +82 | DeepSeek V4 Flash |
| Mystery Game Puzzles (Epoch AI run) | 34% ↗epoch run | 32% ↗epoch run | +2 pt | DeepSeek V4 Flash |
| OTIS Mock AIME 2024-2025 (Epoch AI run) | 94.4% ↗epoch run | 95.6% ↗epoch run | -1.2 pt | Gemini 3.5 Flash |
| SimpleBench (SimpleBench) | 61.1% ↗ | 76.7% ↗ | -15.6 pt | Gemini 3.5 Flash |
| SimpleQA Verified | 33.6% ↗epoch run | 66.2% ↗epoch run | -32.6 pt | Gemini 3.5 Flash |
| Terminal-Bench 4.0 (Vals AI) | 18.7% ↗ | 6.1% ↗ | +12.6 pt | DeepSeek V4 Flash |
| Toolathlon-Verified (HKUST) | 70.7% ↗ | 67.3% ↗ | +3.4 pt | DeepSeek V4 Flash |
| WeirdML (Håvard Tveit Ihle) | 63% ↗ | 62.6% ↗ | +0.4 pt | DeepSeek V4 Flash |
| Only DeepSeek V4 Flash reports · 7 | ||||
| Agents' Last Exam | 25.2% ↗ | not reported | — | — |
| AutomationBench Public | 25.1% ↗ | not reported | — | — |
| Cybergym | 76.7% ↗ | not reported | — | — |
| DeepSWE | 54.4% ↗ | not reported | — | — |
| NL2Repo | 54.2% ↗ | not reported | — | — |
| Terminal Bench 2.1 | 82.7% ↗ | not reported | — | — |
| Toolathlon-Verified | 70.3% ↗ | not reported | — | — |
| Only Gemini 3.5 Flash reports · 15 | ||||
| AIME 2026 | not reported | 95% ↗matharena ⚠ | — | — |
| APEX-Agents (Mercor) | not reported | 27.5% ↗ | — | — |
| ARC-AGI-2 | not reported | 72.1% ↗ | — | — |
| EBR-bench (Epoch AI run) | not reported | 4.8% ↗epoch run | — | — |
| GDPVal-AA | not reported | 1656 ↗ | — | — |
| Humanity's Last Exam (full set, text + MM) | not reported | 40.2% ↗ | — | — |
| LMArena Vision (LMArena) | not reported | 1284.4 ↗ | — | — |
| MCP Atlas | not reported | 83.6% ↗ | — | — |
| MMMU-Pro | not reported | 83.6% ↗ | — | — |
| OSWorld-Verified | not reported | 78.4% ↗ | — | — |
| SAGE (Vals AI) | not reported | 49.9% ↗ | — | — |
| SWE-Bench Pro (Public) | not reported | 55.1% ↗ | — | — |
| SWE-bench Verified (Epoch AI run) | not reported | 79.3% ↗epoch run | — | — |
| Terminal-bench 2.1 | not reported | 76.2% ↗ | — | — |
| Vending-Bench 2 (Andon Labs) | not reported | 5396.42 ↗ | — | — |
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is DeepSeek V4 Flash minus Gemini 3.5 Flash in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing
Side by side
Official pricing pages and model cards| Spec | DeepSeek V4 Flash | Gemini 3.5 Flash | Edge |
|---|---|---|---|
| Context window | 1M tokens | 1.0M tokens | Gemini 3.5 Flash |
| Max output | — | 66K tokens | — |
| Input price / 1M | — | $1.5 ↗ | — |
| Output price / 1M | — | $9 ↗ | — |
| Cached input / 1M | — | $0.15 ↗ | — |
| Input + output / 1M Lower is cheaper. List prices; batch, tool and regional fees excluded. | — | $10.5 | — |
| Modalities | text | text · vision · audio | Gemini 3.5 Flash |
| Released | 2026-07-31 | 2026-05-19 | — |
| Cited benchmark scores | 28 | 42 | — |
Reliability
Provider status
Live from /statusMore matchups
DeepSeek V4 Flash vs …
Models sharing the most benchmarksMore matchups
Gemini 3.5 Flash vs …
Models sharing the most benchmarksBuilt by Respan
Which one wins on your data?
Public benchmarks are a starting point. Run DeepSeek V4 Flash and Gemini 3.5 Flash on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.