Compare / head-to-head
Gemini 3.8 Flashvs
Qwen3.8-Max
Gemini 3.8 Flash leads 13 of 24 shared benchmarks. Gemini 3.8 Flash is 1.8x cheaper per token. Gemini 3.8 Flash has the larger context window (1.0M tokens).
Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.
Shared benchmarks
13 – 10
Gemini 3.8 Flash leads · 1 tied
Cheaper per token
Gemini 3.8 Flash
1.8x cheaper, input + output
Larger context
Gemini 3.8 Flash
1.0M tokens
Providers
2 providers
Google · Qwen
Gemini 3.8 Flash
Google · released 2026-09-02
textvisionaudio
Qwen3.8-Max
Qwen · released 2026-08-02
textvision
Quality
Benchmark matrix
24 shared · 13 only Gemini 3.8 Flash · 3 only Qwen3.8-Max| Benchmark | Gemini 3.8 Flash | Qwen3.8-Max | Δ | Edge |
|---|---|---|---|---|
| Reported by both · 24 | ||||
| GPQA Diamond | 95.3% ↗ | 92.6% ↗ | +2.7 pt | Gemini 3.8 Flash |
| LMArena Elo | 1494.8 ↗ | 1480.6 ↗ | +14.2 | Gemini 3.8 Flash |
| APEX-Agents (Mercor) | 64.3% ↗ | 63.3% ↗ | +1 pt | Gemini 3.8 Flash |
| Chess Puzzles (Epoch AI run) | 61% ↗epoch run | 40% ↗epoch run | +21 pt | Gemini 3.8 Flash |
| DeepSWE v1.1 (Datacurve) | 73.8% ↗ | 57.5% ↗ | +16.3 pt | Gemini 3.8 Flash |
| FrontierMath Tier 4 v2 (Epoch AI run) | 22% ↗epoch run | 46.3% ↗epoch run | -24.3 pt | Qwen3.8-Max |
| FrontierMath Tiers 1-3 v2 (Epoch AI run) | 68.4% ↗epoch run | 74.7% ↗epoch run | -6.3 pt | Qwen3.8-Max |
| FrontierSWE V2 (Proximal Labs) | 19.6% ↗ | 17.8% ↗ | +1.8 pt | Gemini 3.8 Flash |
| Furniture Assembly (Epoch AI run) | 31.7% ↗epoch run | 20% ↗epoch run | +11.7 pt | Gemini 3.8 Flash |
| LiveBench Agentic Coding (LiveBench) | 54.2% ↗ | 64.6% ↗ | -10.4 pt | Qwen3.8-Max |
| LiveBench Coding (LiveBench) | 72.5% ↗ | 72.9% ↗ | -0.4 pt | Qwen3.8-Max |
| LiveBench Data Analysis (LiveBench) | 54% ↗ | 78.4% ↗ | -24.4 pt | Qwen3.8-Max |
| LiveBench Instruction Following (LiveBench) | 81.4% ↗ | 74.1% ↗ | +7.3 pt | Gemini 3.8 Flash |
| LiveBench Language (LiveBench) | 87.8% ↗ | 79.7% ↗ | +8.1 pt | Gemini 3.8 Flash |
| LiveBench Mathematics (LiveBench) | 91.6% ↗ | 91.3% ↗ | +0.3 pt | Gemini 3.8 Flash |
| LiveBench Reasoning (LiveBench) | 89.3% ↗ | 88.2% ↗ | +1.1 pt | Gemini 3.8 Flash |
| LMArena Agent (LMArena) | 0.0297 ↗ | 0.0251 ↗ | 0 | Tie |
| LMArena Vision (LMArena) | 1290.3 ↗ | 1301.2 ↗ | -10.9 | Qwen3.8-Max |
| LMArena WebDev (LMArena) | 1582.7 ↗ | 1671.2 ↗ | -88.5 | Qwen3.8-Max |
| Mystery Game Puzzles (Epoch AI run) | 47% ↗epoch run | 38% ↗epoch run | +9 pt | Gemini 3.8 Flash |
| OTIS Mock AIME 2024-2025 (Epoch AI run) | 98.9% ↗epoch run | 99.4% ↗epoch run | -0.5 pt | Qwen3.8-Max |
| SAGE (Vals AI) | 35.1% ↗ | 51.3% ↗ | -16.2 pt | Qwen3.8-Max |
| SimpleQA Verified | 69.7% ↗epoch run | 45.8% ↗epoch run | +23.9 pt | Gemini 3.8 Flash |
| Terminal-Bench 4.0 (Vals AI) | 19.2% ↗ | 34.3% ↗ | -15.1 pt | Qwen3.8-Max |
| Only Gemini 3.8 Flash reports · 13 | ||||
| CharXiv Reasoning (no tools) | 86.2% ↗ | not reported | — | — |
| DeepSWE v1.1 | 73.7% ↗ | not reported | — | — |
| GDPVal-AA v2 | 1545 ↗ | not reported | — | — |
| Harvey's Legal Agent Benchmark | 10% ↗ | not reported | — | — |
| HLE-Verified | 54.9% ↗ | not reported | — | — |
| LABBench2 | 86.2% ↗ | not reported | — | — |
| OSWorld-2.0 | 59% ↗ | not reported | — | — |
| SimpleBench (SimpleBench) | 82.4% ↗ | not reported | — | — |
| Terminal-bench 2.1 | 89.4% ↗ | not reported | — | — |
| Terminal-bench 4.0 | 19.1% ↗ | not reported | — | — |
| Vals Finance Agent v2 | 61.4% ↗ | not reported | — | — |
| Vending-Bench 2 (Andon Labs) | 5093.79 ↗ | not reported | — | — |
| WeirdML (Håvard Tveit Ihle) | 84.8% ↗ | not reported | — | — |
| Only Qwen3.8-Max reports · 3 | ||||
| SWE-Bench Pro | not reported | 67.7% ↗ | — | — |
| tau2-bench Banking Knowledge (Sierra) | not reported | 55.2% ↗ | — | — |
| Terminal-Bench 2.1 | not reported | 86.6% ↗ | — | — |
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Gemini 3.8 Flash minus Qwen3.8-Max in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing
Side by side
Official pricing pages and model cards| Spec | Gemini 3.8 Flash | Qwen3.8-Max | Edge |
|---|---|---|---|
| Context window | 1.0M tokens | 1M tokens | Gemini 3.8 Flash |
| Max output | 66K tokens | 131K tokens | Qwen3.8-Max |
| Input price / 1M | $0.75 ↗ | $2 ↗ | Gemini 3.8 Flash |
| Output price / 1M | $3.75 ↗ | $6 ↗ | Gemini 3.8 Flash |
| Cached input / 1M | $0.075 ↗ | — | — |
| Input + output / 1M Lower is cheaper. List prices; batch, tool and regional fees excluded. | $4.5 | $8 | Gemini 3.8 Flash |
| Modalities | text · vision · audio | text · vision | Gemini 3.8 Flash |
| Released | 2026-09-02 | 2026-08-02 | — |
| Cited benchmark scores | 43 | 34 | — |
Reliability
Provider status
Live from /statusMore matchups
Gemini 3.8 Flash vs …
Models sharing the most benchmarksMore matchups
Qwen3.8-Max vs …
Models sharing the most benchmarksBuilt by Respan
Which one wins on your data?
Public benchmarks are a starting point. Run Gemini 3.8 Flash and Qwen3.8-Max on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.