Compare / head-to-head

Claude Opus 5.5vsGemini 3.1 Pro Preview

Claude Opus 5.5 leads 20 of 24 shared benchmarks. Gemini 3.1 Pro Preview is 1.7x cheaper per token. Gemini 3.1 Pro Preview has the larger context window (1.0M tokens).

Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.

Shared benchmarks
20 – 4
Claude Opus 5.5 leads
Cheaper per token
Gemini 3.1 Pro Preview
1.7x cheaper, input + output
Larger context
Gemini 3.1 Pro Preview
1.0M tokens
Providers
2 providers
Anthropic · Google
Claude Opus 5.5
Anthropic · released 2026-09-22
textvision
Context
1M
Max out
128K
Input /1M
$4 ↗
Output /1M
$20 ↗
Cached /1M
$0.2
Scores
34 · 9 core
Gemini 3.1 Pro Preview
Google · released 2026-02-19
textvisionaudio
Context
1.0M
Max out
66K
Input /1M
$2 ↗
Output /1M
$12 ↗
Cached /1M
$0.2
Scores
43 · 12 core
Quality

Benchmark matrix

BenchmarkClaude Opus 5.5Gemini 3.1 Pro PreviewΔEdge
Reported by both · 24
GPQA Diamond90.6% ↗epoch run94.3% ↗-3.7 ptGemini 3.1 Pro Preview
LMArena Elo1503.7 ↗1480.1 ↗+23.6Claude Opus 5.5
APEX-Agents (Mercor)73.5% ↗35.3% ↗+38.2 ptClaude Opus 5.5
EBR-bench (Epoch AI run)71.4% ↗epoch run14.3% ↗epoch run+57.1 ptClaude Opus 5.5
FrontierMath Tier 4 v2 (Epoch AI run)95% ↗epoch run26.8% ↗epoch run+68.2 ptClaude Opus 5.5
FrontierMath Tiers 1-3 v2 (Epoch AI run)91.2% ↗epoch run59.6% ↗epoch run+31.6 ptClaude Opus 5.5
Furniture Assembly (Epoch AI run)83.3% ↗epoch run26.7% ↗epoch run+56.6 ptClaude Opus 5.5
LiveBench Agentic Coding (LiveBench)71.7% ↗44.1% ↗+27.6 ptClaude Opus 5.5
LiveBench Coding (LiveBench)89.3% ↗76.5% ↗+12.8 ptClaude Opus 5.5
LiveBench Data Analysis (LiveBench)80.3% ↗78.5% ↗+1.8 ptClaude Opus 5.5
LiveBench Instruction Following (LiveBench)67% ↗79.1% ↗-12.1 ptGemini 3.1 Pro Preview
LiveBench Language (LiveBench)86.3% ↗85.4% ↗+0.9 ptClaude Opus 5.5
LiveBench Mathematics (LiveBench)97.1% ↗91% ↗+6.1 ptClaude Opus 5.5
LiveBench Reasoning (LiveBench)92.2% ↗84% ↗+8.2 ptClaude Opus 5.5
LMArena Agent (LMArena)0.1382 ↗-0.0771 ↗+0.2Claude Opus 5.5
LMArena WebDev (LMArena)1815.4 ↗1446.2 ↗+369.2Claude Opus 5.5
MirrorCode (Epoch AI run)77.4% ↗epoch run8.9% ↗epoch run+68.5 ptClaude Opus 5.5
Mystery Game Puzzles (Epoch AI run)71% ↗epoch run34% ↗epoch run+37 ptClaude Opus 5.5
OTIS Mock AIME 2024-2025 (Epoch AI run)100% ↗epoch run95.6% ↗epoch run+4.4 ptClaude Opus 5.5
SAGE (Vals AI)45.8% ↗48.7% ↗-2.9 ptGemini 3.1 Pro Preview
SimpleBench (SimpleBench)88.4% ↗79.6% ↗+8.8 ptClaude Opus 5.5
SimpleQA Verified72.2% ↗epoch run73.5% ↗epoch run-1.3 ptGemini 3.1 Pro Preview
Terminal-Bench 4.0 (Vals AI)65.2% ↗2.5% ↗+62.7 ptClaude Opus 5.5
Vending-Bench 2 (Andon Labs)9235.25 ↗3774.25 ↗+5461Claude Opus 5.5
Only Claude Opus 5.5 reports · 10
AutomationBench 1.0.640% ↗not reported——
Chartography (with tools)89% ↗not reported——
CursorBench 4.057.8% ↗not reported——
FrontierCode v1.1 (Main)54.4% ↗not reported——
FrontierSWE V2 (Proximal Labs)62.3% ↗not reported——
GDPval-AA v2.11846 ↗not reported——
Humanity's Last Exam (with tools)67.7% ↗not reported——
OSWorld 2.0 (partial)81.8% ↗not reported——
Terminal-Bench 4.066.4% ↗not reported——
Terminal-Bench Science 0.158.7% ↗not reported——
Only Gemini 3.1 Pro Preview reports · 12
AIME 2026not reported98.3% ↗matharena ⚠——
ARC-AGI-2not reported77.1% ↗——
BALROG (BALROG)not reported57% ↗——
Chess Puzzles (Epoch AI run)not reported55% ↗epoch run——
DeepSWE v1.1 (Datacurve)not reported11.7% ↗——
GSO Opt@1 (GSO)not reported21.6% ↗——
LMArena Vision (LMArena)not reported1279.3 ↗——
SWE-Bench Pronot reported54.2% ↗——
SWE-bench Verified (Epoch AI run)not reported75.6% ↗epoch run——
tau2-bench Banking Knowledge (Sierra)not reported26% ↗——
Toolathlon-Verified (HKUST)not reported61.1% ↗——
WeirdML (Håvard Tveit Ihle)not reported72.1% ↗——
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Claude Opus 5.5 minus Gemini 3.1 Pro Preview in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing

Side by side

SpecClaude Opus 5.5Gemini 3.1 Pro PreviewEdge
Context window1M tokens1.0M tokensGemini 3.1 Pro Preview
Max output128K tokens66K tokensClaude Opus 5.5
Input price / 1M$4 ↗$2 ↗Gemini 3.1 Pro Preview
Output price / 1M$20 ↗$12 ↗Gemini 3.1 Pro Preview
Cached input / 1M$0.2 ↗$0.2 ↗Tie
Input + output / 1M
Lower is cheaper. List prices; batch, tool and regional fees excluded.
$24$14Gemini 3.1 Pro Preview
Modalitiestext · visiontext · vision · audioGemini 3.1 Pro Preview
Released2026-09-222026-02-19—
Cited benchmark scores3443—
Reliability

Provider status

All providers →
More matchups

Claude Opus 5.5 vs …

More matchups

Gemini 3.1 Pro Preview vs …

Built by Respan
Which one wins on your data?

Public benchmarks are a starting point. Run Claude Opus 5.5 and Gemini 3.1 Pro Preview on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.