Compare / head-to-head

Claude Fable 5vsGemini 3.1 Pro Preview

Claude Fable 5 leads 27 of 31 shared benchmarks. Gemini 3.1 Pro Preview is 4.3x cheaper per token. Gemini 3.1 Pro Preview has the larger context window (1.0M tokens).

Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.

Shared benchmarks
27 – 4
Claude Fable 5 leads
Cheaper per token
Gemini 3.1 Pro Preview
4.3x cheaper, input + output
Larger context
Gemini 3.1 Pro Preview
1.0M tokens
Providers
2 providers
Anthropic · Google
Claude Fable 5
Anthropic · released 2026-06-09
textvision
Context
1M
Max out
128K
Input /1M
$10 ↗
Output /1M
$50 ↗
Cached /1M
$1
Scores
44 · 11 core
Gemini 3.1 Pro Preview
Google · released 2026-02-19
textvisionaudio
Context
1.0M
Max out
66K
Input /1M
$2 ↗
Output /1M
$12 ↗
Cached /1M
$0.2
Scores
43 · 12 core
Quality

Benchmark matrix

BenchmarkClaude Fable 5Gemini 3.1 Pro PreviewΔEdge
Reported by both · 31
GPQA Diamond85.9% ↗epoch run94.3% ↗-8.4 ptGemini 3.1 Pro Preview
LMArena Elo1504.3 ↗1480.1 ↗+24.2Claude Fable 5
APEX-Agents (Mercor)63.6% ↗35.3% ↗+28.3 ptClaude Fable 5
Chess Puzzles (Epoch AI run)41% ↗epoch run55% ↗epoch run-14 ptGemini 3.1 Pro Preview
DeepSWE v1.1 (Datacurve)69.9% ↗11.7% ↗+58.2 ptClaude Fable 5
EBR-bench (Epoch AI run)39.5% ↗epoch run14.3% ↗epoch run+25.2 ptClaude Fable 5
FrontierMath Tier 4 v2 (Epoch AI run)90.2% ↗epoch run26.8% ↗epoch run+63.4 ptClaude Fable 5
FrontierMath Tiers 1-3 v2 (Epoch AI run)87% ↗epoch run59.6% ↗epoch run+27.4 ptClaude Fable 5
Furniture Assembly (Epoch AI run)35.8% ↗epoch run26.7% ↗epoch run+9.1 ptClaude Fable 5
GSO Opt@1 (GSO)76.5% ↗21.6% ↗+54.9 ptClaude Fable 5
LiveBench Agentic Coding (LiveBench)62.2% ↗44.1% ↗+18.1 ptClaude Fable 5
LiveBench Coding (LiveBench)86% ↗76.5% ↗+9.5 ptClaude Fable 5
LiveBench Data Analysis (LiveBench)80.5% ↗78.5% ↗+2 ptClaude Fable 5
LiveBench Instruction Following (LiveBench)75.8% ↗79.1% ↗-3.3 ptGemini 3.1 Pro Preview
LiveBench Language (LiveBench)90.7% ↗85.4% ↗+5.3 ptClaude Fable 5
LiveBench Mathematics (LiveBench)96% ↗91% ↗+5 ptClaude Fable 5
LiveBench Reasoning (LiveBench)89.7% ↗84% ↗+5.7 ptClaude Fable 5
LMArena Agent (LMArena)0.0821 ↗-0.0771 ↗+0.2Claude Fable 5
LMArena Vision (LMArena)1308.8 ↗1279.3 ↗+29.5Claude Fable 5
LMArena WebDev (LMArena)1625.3 ↗1446.2 ↗+179.1Claude Fable 5
MirrorCode (Epoch AI run)63.9% ↗epoch run8.9% ↗epoch run+55 ptClaude Fable 5
Mystery Game Puzzles (Epoch AI run)52% ↗epoch run34% ↗epoch run+18 ptClaude Fable 5
OTIS Mock AIME 2024-2025 (Epoch AI run)100% ↗epoch run95.6% ↗epoch run+4.4 ptClaude Fable 5
SAGE (Vals AI)51.9% ↗48.7% ↗+3.2 ptClaude Fable 5
SimpleBench (SimpleBench)81.9% ↗79.6% ↗+2.3 ptClaude Fable 5
SimpleQA Verified70.7% ↗epoch run73.5% ↗epoch run-2.8 ptGemini 3.1 Pro Preview
SWE-Bench Pro80% ↗54.2% ↗+25.8 ptClaude Fable 5
tau2-bench Banking Knowledge (Sierra)39.7% ↗26% ↗+13.7 ptClaude Fable 5
Terminal-Bench 4.0 (Vals AI)41.4% ↗2.5% ↗+38.9 ptClaude Fable 5
Vending-Bench 2 (Andon Labs)5680.26 ↗3774.25 ↗+1906Claude Fable 5
WeirdML (Håvard Tveit Ihle)91.9% ↗72.1% ↗+19.8 ptClaude Fable 5
Only Claude Fable 5 reports · 13
SWE-bench Verified95% ↗not reported——
AutomationBench17.4% ↗not reported——
Blueprint-Bench 238.6% ↗not reported——
CursorBench72.9% ↗not reported——
FrontierCode (Diamond)29.3% ↗not reported——
FrontierSWE V2 (Proximal Labs)47% ↗not reported——
GDP.pdf29.8% ↗not reported——
Legal Agent Benchmark (Harvey's Held-Out Set)13.3% ↗not reported——
OfficeQA Pro57.9% ↗not reported——
OSWorld-Verified85% ↗not reported——
OSWorld-Verified (XLANG)86% ↗not reported——
SWE-rebench 2026-05-15 to 2026-07-01 (Nebius)64.5% ↗not reported——
Terminal-Bench 2.184.3% ↗not reported——
Only Gemini 3.1 Pro Preview reports · 5
AIME 2026not reported98.3% ↗matharena ⚠——
ARC-AGI-2not reported77.1% ↗——
BALROG (BALROG)not reported57% ↗——
SWE-bench Verified (Epoch AI run)not reported75.6% ↗epoch run——
Toolathlon-Verified (HKUST)not reported61.1% ↗——
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Claude Fable 5 minus Gemini 3.1 Pro Preview in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing

Side by side

SpecClaude Fable 5Gemini 3.1 Pro PreviewEdge
Context window1M tokens1.0M tokensGemini 3.1 Pro Preview
Max output128K tokens66K tokensClaude Fable 5
Input price / 1M$10 ↗$2 ↗Gemini 3.1 Pro Preview
Output price / 1M$50 ↗$12 ↗Gemini 3.1 Pro Preview
Cached input / 1M$1 ↗$0.2 ↗Gemini 3.1 Pro Preview
Input + output / 1M
Lower is cheaper. List prices; batch, tool and regional fees excluded.
$60$14Gemini 3.1 Pro Preview
Modalitiestext · visiontext · vision · audioGemini 3.1 Pro Preview
Released2026-06-092026-02-19—
Cited benchmark scores4443—
Reliability

Provider status

All providers →
More matchups

Claude Fable 5 vs …

More matchups

Gemini 3.1 Pro Preview vs …

Built by Respan
Which one wins on your data?

Public benchmarks are a starting point. Run Claude Fable 5 and Gemini 3.1 Pro Preview on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.