Compare / head-to-head

Gemini 3.8 FlashvsQwen3.8-Max

Gemini 3.8 Flash leads 13 of 24 shared benchmarks. Gemini 3.8 Flash is 1.8x cheaper per token. Gemini 3.8 Flash has the larger context window (1.0M tokens).

Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.

Shared benchmarks
13 – 10
Gemini 3.8 Flash leads · 1 tied
Cheaper per token
Gemini 3.8 Flash
1.8x cheaper, input + output
Larger context
Gemini 3.8 Flash
1.0M tokens
Providers
2 providers
Google · Qwen
Gemini 3.8 Flash
Google · released 2026-09-02
textvisionaudio
Context
1.0M
Max out
66K
Input /1M
$0.75 ↗
Output /1M
$3.75 ↗
Cached /1M
$0.075
Scores
43 · 10 core
Qwen3.8-Max
Qwen · released 2026-08-02
textvision
Context
1M
Max out
131K
Input /1M
$2 ↗
Output /1M
$6 ↗
Cached /1M
—
Scores
34 · 10 core
Quality

Benchmark matrix

BenchmarkGemini 3.8 FlashQwen3.8-MaxΔEdge
Reported by both · 24
GPQA Diamond95.3% ↗92.6% ↗+2.7 ptGemini 3.8 Flash
LMArena Elo1494.8 ↗1480.6 ↗+14.2Gemini 3.8 Flash
APEX-Agents (Mercor)64.3% ↗63.3% ↗+1 ptGemini 3.8 Flash
Chess Puzzles (Epoch AI run)61% ↗epoch run40% ↗epoch run+21 ptGemini 3.8 Flash
DeepSWE v1.1 (Datacurve)73.8% ↗57.5% ↗+16.3 ptGemini 3.8 Flash
FrontierMath Tier 4 v2 (Epoch AI run)22% ↗epoch run46.3% ↗epoch run-24.3 ptQwen3.8-Max
FrontierMath Tiers 1-3 v2 (Epoch AI run)68.4% ↗epoch run74.7% ↗epoch run-6.3 ptQwen3.8-Max
FrontierSWE V2 (Proximal Labs)19.6% ↗17.8% ↗+1.8 ptGemini 3.8 Flash
Furniture Assembly (Epoch AI run)31.7% ↗epoch run20% ↗epoch run+11.7 ptGemini 3.8 Flash
LiveBench Agentic Coding (LiveBench)54.2% ↗64.6% ↗-10.4 ptQwen3.8-Max
LiveBench Coding (LiveBench)72.5% ↗72.9% ↗-0.4 ptQwen3.8-Max
LiveBench Data Analysis (LiveBench)54% ↗78.4% ↗-24.4 ptQwen3.8-Max
LiveBench Instruction Following (LiveBench)81.4% ↗74.1% ↗+7.3 ptGemini 3.8 Flash
LiveBench Language (LiveBench)87.8% ↗79.7% ↗+8.1 ptGemini 3.8 Flash
LiveBench Mathematics (LiveBench)91.6% ↗91.3% ↗+0.3 ptGemini 3.8 Flash
LiveBench Reasoning (LiveBench)89.3% ↗88.2% ↗+1.1 ptGemini 3.8 Flash
LMArena Agent (LMArena)0.0297 ↗0.0251 ↗0Tie
LMArena Vision (LMArena)1290.3 ↗1301.2 ↗-10.9Qwen3.8-Max
LMArena WebDev (LMArena)1582.7 ↗1671.2 ↗-88.5Qwen3.8-Max
Mystery Game Puzzles (Epoch AI run)47% ↗epoch run38% ↗epoch run+9 ptGemini 3.8 Flash
OTIS Mock AIME 2024-2025 (Epoch AI run)98.9% ↗epoch run99.4% ↗epoch run-0.5 ptQwen3.8-Max
SAGE (Vals AI)35.1% ↗51.3% ↗-16.2 ptQwen3.8-Max
SimpleQA Verified69.7% ↗epoch run45.8% ↗epoch run+23.9 ptGemini 3.8 Flash
Terminal-Bench 4.0 (Vals AI)19.2% ↗34.3% ↗-15.1 ptQwen3.8-Max
Only Gemini 3.8 Flash reports · 13
CharXiv Reasoning (no tools)86.2% ↗not reported——
DeepSWE v1.173.7% ↗not reported——
GDPVal-AA v21545 ↗not reported——
Harvey's Legal Agent Benchmark10% ↗not reported——
HLE-Verified54.9% ↗not reported——
LABBench286.2% ↗not reported——
OSWorld-2.059% ↗not reported——
SimpleBench (SimpleBench)82.4% ↗not reported——
Terminal-bench 2.189.4% ↗not reported——
Terminal-bench 4.019.1% ↗not reported——
Vals Finance Agent v261.4% ↗not reported——
Vending-Bench 2 (Andon Labs)5093.79 ↗not reported——
WeirdML (Håvard Tveit Ihle)84.8% ↗not reported——
Only Qwen3.8-Max reports · 3
SWE-Bench Pronot reported67.7% ↗——
tau2-bench Banking Knowledge (Sierra)not reported55.2% ↗——
Terminal-Bench 2.1not reported86.6% ↗——
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Gemini 3.8 Flash minus Qwen3.8-Max in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing

Side by side

SpecGemini 3.8 FlashQwen3.8-MaxEdge
Context window1.0M tokens1M tokensGemini 3.8 Flash
Max output66K tokens131K tokensQwen3.8-Max
Input price / 1M$0.75 ↗$2 ↗Gemini 3.8 Flash
Output price / 1M$3.75 ↗$6 ↗Gemini 3.8 Flash
Cached input / 1M$0.075 ↗——
Input + output / 1M
Lower is cheaper. List prices; batch, tool and regional fees excluded.
$4.5$8Gemini 3.8 Flash
Modalitiestext · vision · audiotext · visionGemini 3.8 Flash
Released2026-09-022026-08-02—
Cited benchmark scores4334—
Reliability

Provider status

All providers →
More matchups

Gemini 3.8 Flash vs …

Built by Respan
Which one wins on your data?

Public benchmarks are a starting point. Run Gemini 3.8 Flash and Qwen3.8-Max on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.