Compare / head-to-head

DeepSeek V3.2vsQwen3.5 397B-A17B

Qwen3.5 397B-A17B leads 6 of 7 shared benchmarks. Qwen3.5 397B-A17B has the larger context window (262K tokens).

Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.

Shared benchmarks
0 – 6
Qwen3.5 397B-A17B leads · 1 tied
Cheaper per token
—
list price, input + output
Larger context
Qwen3.5 397B-A17B
262K tokens
Providers
2 providers
DeepSeek · Qwen
DeepSeek V3.2
DeepSeek · released 2025-12-01
text
Context
131K
Max out
66K
Input /1M
—
Output /1M
—
Cached /1M
—
Scores
11 · 8 core
Qwen3.5 397B-A17B
Qwen · released 2026-02-15
textvision
Context
262K
Max out
—
Input /1M
$0.6 ↗
Output /1M
$3.6 ↗
Cached /1M
—
Scores
22 · 9 core
Quality

Benchmark matrix

BenchmarkDeepSeek V3.2Qwen3.5 397B-A17BΔEdge
Reported by both · 7
SWE-bench Verified73.1% ↗76.4% ↗-3.3 ptQwen3.5 397B-A17B
GPQA Diamond82.4% ↗88.4% ↗-6 ptQwen3.5 397B-A17B
MMLU-Pro85% ↗87.8% ↗-2.8 ptQwen3.5 397B-A17B
AIME 202694.2% ↗matharena94.2% ↗matharena ⚠0Tie
LMArena Elo1424.8 ↗1438.3 ↗-13.5Qwen3.5 397B-A17B
APEX-Agents (Mercor)21.3% ↗24.9% ↗-3.6 ptQwen3.5 397B-A17B
LMArena WebDev (LMArena)1361.7 ↗1399.4 ↗-37.7Qwen3.5 397B-A17B
Only DeepSeek V3.2 reports · 2
AIME 202593.1% ↗not reported——
Humanity's Last Exam (no tools)25.1% ↗not reported——
Only Qwen3.5 397B-A17B reports · 10
Chess Puzzles (Epoch AI run)not reported13% ↗epoch run——
FrontierMath Tiers 1-3 v2 (Epoch AI run)not reported29.5% ↗epoch run——
LMArena Vision (LMArena)not reported1246.9 ↗——
Mystery Game Puzzles (Epoch AI run)not reported18% ↗epoch run——
OTIS Mock AIME 2024-2025 (Epoch AI run)not reported88.9% ↗epoch run——
tau2-bench Airline (Sierra)not reported81.5% ↗——
tau2-bench Banking Knowledge (Sierra)not reported9.8% ↗——
tau2-bench Retail (Sierra)not reported84.4% ↗——
tau2-bench Telecom (Sierra)not reported97.8% ↗——
Toolathlon-Verified (HKUST)not reported40.7% ↗——
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is DeepSeek V3.2 minus Qwen3.5 397B-A17B in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing

Side by side

SpecDeepSeek V3.2Qwen3.5 397B-A17BEdge
Context window131K tokens262K tokensQwen3.5 397B-A17B
Max output66K tokens——
Input price / 1M—$0.6 ↗—
Output price / 1M—$3.6 ↗—
Cached input / 1M———
Input + output / 1M
Lower is cheaper. List prices; batch, tool and regional fees excluded.
—$4.2—
Modalitiestexttext · visionQwen3.5 397B-A17B
Released2025-12-012026-02-15—
Cited benchmark scores1122—
Reliability

Provider status

All providers →
More matchups

DeepSeek V3.2 vs …

Built by Respan
Which one wins on your data?

Public benchmarks are a starting point. Run DeepSeek V3.2 and Qwen3.5 397B-A17B on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.