Compare / head-to-head

Grok 4.6vsGrok 4.7

Grok 4.6 leads 12 of 24 shared benchmarks. Both list the same combined token price. Both offer a 500K-token context window.

Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.

Shared benchmarks
12 – 11
Grok 4.6 leads · 1 tied
Cheaper per token
Tie
list price, input + output
Larger context
Tie
both 500K tokens
Providers
xAI
same provider
Grok 4.6
xAI · released 2026-08-12
textvision
Context
500K
Max out
—
Input /1M
$2 ↗
Output /1M
$6 ↗
Cached /1M
$0.5
Scores
43 · 9 core
Grok 4.7
xAI · released 2026-09-21
textvision
Context
500K
Max out
—
Input /1M
$2 ↗
Output /1M
$6 ↗
Cached /1M
$0.5
Scores
30 · 8 core
Quality

Benchmark matrix

BenchmarkGrok 4.6Grok 4.7ΔEdge
Reported by both · 24
GPQA Diamond94% ↗epoch run92.7% ↗epoch run+1.3 ptGrok 4.6
LMArena Elo1453.8 ↗1442.5 ↗+11.3Grok 4.6
APEX-Agents (Mercor)65.3% ↗54.6% ↗+10.7 ptGrok 4.6
Chess Puzzles (Epoch AI run)40% ↗epoch run38% ↗epoch run+2 ptGrok 4.6
DeepSWE v1.165.9% ↗71% ↗-5.1 ptGrok 4.7
FrontierMath Tier 4 v2 (Epoch AI run)31.7% ↗epoch run17.1% ↗epoch run+14.6 ptGrok 4.6
FrontierMath Tiers 1-3 v2 (Epoch AI run)66% ↗epoch run53% ↗epoch run+13 ptGrok 4.6
FrontierSWE V2 (Proximal Labs)25.3% ↗29.5% ↗-4.2 ptGrok 4.7
Furniture Assembly (Epoch AI run)40% ↗epoch run20.8% ↗epoch run+19.2 ptGrok 4.6
LiveBench Agentic Coding (LiveBench)57% ↗54% ↗+3 ptGrok 4.6
LiveBench Coding (LiveBench)76.8% ↗77.2% ↗-0.4 ptGrok 4.7
LiveBench Data Analysis (LiveBench)73.9% ↗76.9% ↗-3 ptGrok 4.7
LiveBench Instruction Following (LiveBench)71.9% ↗75.3% ↗-3.4 ptGrok 4.7
LiveBench Language (LiveBench)83.7% ↗80.1% ↗+3.6 ptGrok 4.6
LiveBench Mathematics (LiveBench)92.6% ↗95.7% ↗-3.1 ptGrok 4.7
LiveBench Reasoning (LiveBench)90.5% ↗82.7% ↗+7.8 ptGrok 4.6
LMArena Agent (LMArena)0.0128 ↗0.0404 ↗0Tie
LMArena WebDev (LMArena)1619.5 ↗1637.5 ↗-18Grok 4.7
Mystery Game Puzzles (Epoch AI run)34% ↗epoch run29% ↗epoch run+5 ptGrok 4.6
OTIS Mock AIME 2024-2025 (Epoch AI run)99.2% ↗epoch run98.1% ↗epoch run+1.1 ptGrok 4.6
SAGE (Vals AI)28.9% ↗40.8% ↗-11.9 ptGrok 4.7
SimpleQA Verified49.3% ↗epoch run56% ↗epoch run-6.7 ptGrok 4.7
Terminal-Bench 4.0 (Vals AI)17.2% ↗28.8% ↗-11.6 ptGrok 4.7
Vending-Bench 2 (Andon Labs)9047.03 ↗10536.83 ↗-1489.8Grok 4.7
Only Grok 4.6 reports · 14
AA-Briefcase1577 ↗not reported——
APEX-Agents57.5% ↗not reported——
APEX-SWE56.4% ↗not reported——
Artificial Analysis Intelligence Index61 ↗not reported——
CursorBench 3.269.9% ↗not reported——
DeepSWE v1.1 (Datacurve)67.5% ↗not reported——
EBR-bench (Epoch AI run)30.5% ↗epoch runnot reported——
FrontierCode v1.1 (Extended)61.3% ↗not reported——
GDPVal-AA v21753 ↗not reported——
Harvey LAB (Vals)15.8% ↗not reported——
LMArena Vision (LMArena)1263.5 ↗not reported——
SimpleBench (SimpleBench)75.9% ↗not reported——
Terminal-Bench 3.026% ↗not reported——
WeirdML (Håvard Tveit Ihle)67.3% ↗not reported——
Only Grok 4.7 reports · 6
AA Briefcase v1.1not reported1657 ↗——
CursorBench 4.0not reported46.3% ↗——
EEBenchnot reported64% ↗——
Harvey Legal Agent Benchmarknot reported19.6% ↗——
HealthBench Professionalnot reported56.7% ↗——
Terminal-Bench 4.0not reported38% ↗——
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Grok 4.6 minus Grok 4.7 in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing

Side by side

SpecGrok 4.6Grok 4.7Edge
Context window500K tokens500K tokensTie
Max output———
Input price / 1M$2 ↗$2 ↗Tie
Output price / 1M$6 ↗$6 ↗Tie
Cached input / 1M$0.5 ↗$0.5 ↗Tie
Input + output / 1M
Lower is cheaper. List prices; batch, tool and regional fees excluded.
$8$8Tie
Modalitiestext · visiontext · visionTie
Released2026-08-122026-09-21—
Cited benchmark scores4330—
Reliability

xAI status

All providers →
More matchups

Grok 4.6 vs …

Built by Respan
Which one wins on your data?

Public benchmarks are a starting point. Run Grok 4.6 and Grok 4.7 on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.