Compare / head-to-head

Claude Sonnet 5.5vsGrok 4.6

Claude Sonnet 5.5 leads 18 of 23 shared benchmarks. Grok 4.6 is 1.5x cheaper per token. Claude Sonnet 5.5 has the larger context window (1M tokens).

Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.

Shared benchmarks
18 – 4
Claude Sonnet 5.5 leads · 1 tied
Cheaper per token
Grok 4.6
1.5x cheaper, input + output
Larger context
Claude Sonnet 5.5
1M tokens
Providers
2 providers
Anthropic · xAI
Claude Sonnet 5.5
Anthropic · released 2026-09-28
textvision
Context
1M
Max out
128K
Input /1M
$2 ↗
Output /1M
$10 ↗
Cached /1M
$0.2
Scores
29 · 10 core
Grok 4.6
xAI · released 2026-08-12
textvision
Context
500K
Max out
—
Input /1M
$2 ↗
Output /1M
$6 ↗
Cached /1M
$0.5
Scores
43 · 9 core
Quality

Benchmark matrix

BenchmarkClaude Sonnet 5.5Grok 4.6ΔEdge
Reported by both · 23
GPQA Diamond95.6% ↗epoch run94% ↗epoch run+1.6 ptClaude Sonnet 5.5
LMArena Elo1471 ↗1453.8 ↗+17.2Claude Sonnet 5.5
APEX-Agents (Mercor)75.5% ↗65.3% ↗+10.2 ptClaude Sonnet 5.5
FrontierMath Tier 4 v2 (Epoch AI run)80.5% ↗epoch run31.7% ↗epoch run+48.8 ptClaude Sonnet 5.5
FrontierMath Tiers 1-3 v2 (Epoch AI run)88.8% ↗epoch run66% ↗epoch run+22.8 ptClaude Sonnet 5.5
FrontierSWE V2 (Proximal Labs)61.9% ↗25.3% ↗+36.6 ptClaude Sonnet 5.5
Furniture Assembly (Epoch AI run)75% ↗epoch run40% ↗epoch run+35 ptClaude Sonnet 5.5
LiveBench Agentic Coding (LiveBench)56.3% ↗57% ↗-0.7 ptGrok 4.6
LiveBench Coding (LiveBench)91.4% ↗76.8% ↗+14.6 ptClaude Sonnet 5.5
LiveBench Data Analysis (LiveBench)78.6% ↗73.9% ↗+4.7 ptClaude Sonnet 5.5
LiveBench Instruction Following (LiveBench)70.5% ↗71.9% ↗-1.4 ptGrok 4.6
LiveBench Language (LiveBench)83.4% ↗83.7% ↗-0.3 ptGrok 4.6
LiveBench Mathematics (LiveBench)96.7% ↗92.6% ↗+4.1 ptClaude Sonnet 5.5
LiveBench Reasoning (LiveBench)91.6% ↗90.5% ↗+1.1 ptClaude Sonnet 5.5
LMArena Agent (LMArena)0.1252 ↗0.0128 ↗+0.1Claude Sonnet 5.5
LMArena Vision (LMArena)1268.3 ↗1263.5 ↗+4.8Claude Sonnet 5.5
LMArena WebDev (LMArena)1786.3 ↗1619.5 ↗+166.8Claude Sonnet 5.5
Mystery Game Puzzles (Epoch AI run)65% ↗epoch run34% ↗epoch run+31 ptClaude Sonnet 5.5
OTIS Mock AIME 2024-2025 (Epoch AI run)100% ↗epoch run99.2% ↗epoch run+0.8 ptClaude Sonnet 5.5
SAGE (Vals AI)51.8% ↗28.9% ↗+22.9 ptClaude Sonnet 5.5
SimpleBench (SimpleBench)75.9% ↗75.9% ↗0Tie
SimpleQA Verified46.5% ↗epoch run49.3% ↗epoch run-2.8 ptGrok 4.6
Terminal-Bench 4.0 (Vals AI)64.1% ↗17.2% ↗+46.9 ptClaude Sonnet 5.5
Only Claude Sonnet 5.5 reports · 6
Chartography (no tools)61.6% ↗not reported——
CursorBench 4.055.5% ↗not reported——
FrontierCode v1.1 (Main)46.2% ↗not reported——
Humanity's Last Exam (with tools)64.5% ↗not reported——
OSWorld 2.1 (partial)80.1% ↗not reported——
Terminal-Bench 4.070.6% ↗not reported——
Only Grok 4.6 reports · 15
AA-Briefcasenot reported1577 ↗——
APEX-Agentsnot reported57.5% ↗——
APEX-SWEnot reported56.4% ↗——
Artificial Analysis Intelligence Indexnot reported61 ↗——
Chess Puzzles (Epoch AI run)not reported40% ↗epoch run——
CursorBench 3.2not reported69.9% ↗——
DeepSWE v1.1not reported65.9% ↗——
DeepSWE v1.1 (Datacurve)not reported67.5% ↗——
EBR-bench (Epoch AI run)not reported30.5% ↗epoch run——
FrontierCode v1.1 (Extended)not reported61.3% ↗——
GDPVal-AA v2not reported1753 ↗——
Harvey LAB (Vals)not reported15.8% ↗——
Terminal-Bench 3.0not reported26% ↗——
Vending-Bench 2 (Andon Labs)not reported9047.03 ↗——
WeirdML (Håvard Tveit Ihle)not reported67.3% ↗——
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Claude Sonnet 5.5 minus Grok 4.6 in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing

Side by side

SpecClaude Sonnet 5.5Grok 4.6Edge
Context window1M tokens500K tokensClaude Sonnet 5.5
Max output128K tokens——
Input price / 1M$2 ↗$2 ↗Tie
Output price / 1M$10 ↗$6 ↗Grok 4.6
Cached input / 1M$0.2 ↗$0.5 ↗Claude Sonnet 5.5
Input + output / 1M
Lower is cheaper. List prices; batch, tool and regional fees excluded.
$12$8Grok 4.6
Modalitiestext · visiontext · visionTie
Released2026-09-282026-08-12—
Cited benchmark scores2943—
Reliability

Provider status

All providers →
More matchups

Claude Sonnet 5.5 vs …

More matchups

Grok 4.6 vs …

Built by Respan
Which one wins on your data?

Public benchmarks are a starting point. Run Claude Sonnet 5.5 and Grok 4.6 on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.