Compare / head-to-head

Qwen3.5 0.8BvsQwen3.5 2B

Qwen3.5 2B leads 14 of 14 shared benchmarks. Both offer a 262K-token context window.

Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.

Shared benchmarks
0 – 14
Qwen3.5 2B leads
Cheaper per token
—
list price, input + output
Larger context
Tie
both 262K tokens
Providers
Qwen
same provider
Qwen3.5 0.8B
Qwen · released 2026-03-02
textvisionvideo
Context
262K
Max out
—
Input /1M
—
Output /1M
—
Cached /1M
—
Scores
14 · 5 core
Qwen3.5 2B
Qwen · released 2026-03-02
textvisionvideo
Context
262K
Max out
—
Input /1M
—
Output /1M
—
Cached /1M
—
Scores
16 · 5 core
Quality

Benchmark matrix

BenchmarkQwen3.5 0.8BQwen3.5 2BΔEdge
Reported by both · 14
BFCL-V4 (thinking)25.3% ↗43.6% ↗-18.3 ptQwen3.5 2B
GPQA (thinking)11.9% ↗51.6% ↗-39.7 ptQwen3.5 2B
IFBench (thinking)21% ↗41.3% ↗-20.3 ptQwen3.5 2B
IFEval (non-thinking)52.1% ↗61.2% ↗-9.1 ptQwen3.5 2B
IFEval (thinking)44% ↗78.6% ↗-34.6 ptQwen3.5 2B
LongBench v2 (thinking)26.1% ↗38.7% ↗-12.6 ptQwen3.5 2B
MMLU-Pro (non-thinking)29.7% ↗55.3% ↗-25.6 ptQwen3.5 2B
MMLU-Pro (thinking)42.3% ↗66.5% ↗-24.2 ptQwen3.5 2B
MMLU-Redux (thinking)59.5% ↗79.6% ↗-20.1 ptQwen3.5 2B
MMMLU (thinking)44.3% ↗63.1% ↗-18.8 ptQwen3.5 2B
MMMU (thinking)49% ↗64.2% ↗-15.2 ptQwen3.5 2B
MMMU-Pro (thinking)31.2% ↗50.3% ↗-19.1 ptQwen3.5 2B
SuperGPQA (thinking)21.3% ↗37.5% ↗-16.2 ptQwen3.5 2B
TAU2-Bench (thinking)11.6% ↗48.8% ↗-37.2 ptQwen3.5 2B
Only Qwen3.5 0.8B reports · 0
Qwen3.5 0.8B reports nothing Qwen3.5 2B does not.
Only Qwen3.5 2B reports · 2
Chess Puzzles (Epoch AI run)not reported0% ↗epoch run——
HMMT Feb 25 (thinking)not reported22.9% ↗——
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Qwen3.5 0.8B minus Qwen3.5 2B in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing

Side by side

SpecQwen3.5 0.8BQwen3.5 2BEdge
Context window262K tokens262K tokensTie
Max output———
Input price / 1M———
Output price / 1M———
Cached input / 1M———
Input + output / 1M
Lower is cheaper. List prices; batch, tool and regional fees excluded.
———
Modalitiestext · vision · videotext · vision · videoTie
Released2026-03-022026-03-02—
Cited benchmark scores1416—
Reliability

Qwen status

All providers →
More matchups

Qwen3.5 0.8B vs …

More matchups

Qwen3.5 2B vs …

Built by Respan
Which one wins on your data?

Public benchmarks are a starting point. Run Qwen3.5 0.8B and Qwen3.5 2B on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.