Compare / head-to-head

GPT-6.1 SolvsGrok 4.6

GPT-6.1 Sol leads 22 of 24 shared benchmarks. Grok 4.6 is 1.5x cheaper per token. GPT-6.1 Sol has the larger context window (1.1M tokens).

Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.

Shared benchmarks
22 – 2
GPT-6.1 Sol leads
Cheaper per token
Grok 4.6
1.5x cheaper, input + output
Larger context
GPT-6.1 Sol
1.1M tokens
Providers
2 providers
OpenAI · xAI
GPT-6.1 Sol
OpenAI · released 2026-09-29
textvision
Context
1.1M
Max out
128K
Input /1M
$2 ↗
Output /1M
$10 ↗
Cached /1M
$0.1
Scores
24 · 9 core
Grok 4.6
xAI · released 2026-08-12
textvision
Context
500K
Max out
—
Input /1M
$2 ↗
Output /1M
$6 ↗
Cached /1M
$0.5
Scores
43 · 9 core
Quality

Benchmark matrix

BenchmarkGPT-6.1 SolGrok 4.6ΔEdge
Reported by both · 24
GPQA Diamond95.4% ↗epoch run94% ↗epoch run+1.4 ptGPT-6.1 Sol
LMArena Elo1483.2 ↗1453.8 ↗+29.4GPT-6.1 Sol
APEX-Agents (Mercor)60% ↗65.3% ↗-5.3 ptGrok 4.6
Chess Puzzles (Epoch AI run)61% ↗epoch run40% ↗epoch run+21 ptGPT-6.1 Sol
EBR-bench (Epoch AI run)54.3% ↗epoch run30.5% ↗epoch run+23.8 ptGPT-6.1 Sol
FrontierMath Tier 4 v2 (Epoch AI run)100% ↗epoch run31.7% ↗epoch run+68.3 ptGPT-6.1 Sol
FrontierMath Tiers 1-3 v2 (Epoch AI run)93.7% ↗epoch run66% ↗epoch run+27.7 ptGPT-6.1 Sol
Furniture Assembly (Epoch AI run)80% ↗epoch run40% ↗epoch run+40 ptGPT-6.1 Sol
LiveBench Agentic Coding (LiveBench)56.8% ↗57% ↗-0.2 ptGrok 4.6
LiveBench Coding (LiveBench)80.7% ↗76.8% ↗+3.9 ptGPT-6.1 Sol
LiveBench Data Analysis (LiveBench)82.7% ↗73.9% ↗+8.8 ptGPT-6.1 Sol
LiveBench Instruction Following (LiveBench)74.2% ↗71.9% ↗+2.3 ptGPT-6.1 Sol
LiveBench Language (LiveBench)90.1% ↗83.7% ↗+6.4 ptGPT-6.1 Sol
LiveBench Mathematics (LiveBench)96.8% ↗92.6% ↗+4.2 ptGPT-6.1 Sol
LiveBench Reasoning (LiveBench)92.6% ↗90.5% ↗+2.1 ptGPT-6.1 Sol
LMArena Agent (LMArena)0.1123 ↗0.0128 ↗+0.1GPT-6.1 Sol
LMArena Vision (LMArena)1290.6 ↗1263.5 ↗+27.1GPT-6.1 Sol
LMArena WebDev (LMArena)1757.8 ↗1619.5 ↗+138.3GPT-6.1 Sol
Mystery Game Puzzles (Epoch AI run)80% ↗epoch run34% ↗epoch run+46 ptGPT-6.1 Sol
OTIS Mock AIME 2024-2025 (Epoch AI run)100% ↗epoch run99.2% ↗epoch run+0.8 ptGPT-6.1 Sol
SAGE (Vals AI)46.5% ↗28.9% ↗+17.6 ptGPT-6.1 Sol
SimpleBench (SimpleBench)82.9% ↗75.9% ↗+7 ptGPT-6.1 Sol
SimpleQA Verified73.9% ↗epoch run49.3% ↗epoch run+24.6 ptGPT-6.1 Sol
Terminal-Bench 4.0 (Vals AI)55.1% ↗17.2% ↗+37.9 ptGPT-6.1 Sol
Only GPT-6.1 Sol reports · 0
GPT-6.1 Sol reports nothing Grok 4.6 does not.
Only Grok 4.6 reports · 14
AA-Briefcasenot reported1577 ↗——
APEX-Agentsnot reported57.5% ↗——
APEX-SWEnot reported56.4% ↗——
Artificial Analysis Intelligence Indexnot reported61 ↗——
CursorBench 3.2not reported69.9% ↗——
DeepSWE v1.1not reported65.9% ↗——
DeepSWE v1.1 (Datacurve)not reported67.5% ↗——
FrontierCode v1.1 (Extended)not reported61.3% ↗——
FrontierSWE V2 (Proximal Labs)not reported25.3% ↗——
GDPVal-AA v2not reported1753 ↗——
Harvey LAB (Vals)not reported15.8% ↗——
Terminal-Bench 3.0not reported26% ↗——
Vending-Bench 2 (Andon Labs)not reported9047.03 ↗——
WeirdML (Håvard Tveit Ihle)not reported67.3% ↗——
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is GPT-6.1 Sol minus Grok 4.6 in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing

Side by side

SpecGPT-6.1 SolGrok 4.6Edge
Context window1.1M tokens500K tokensGPT-6.1 Sol
Max output128K tokens——
Input price / 1M$2 ↗$2 ↗Tie
Output price / 1M$10 ↗$6 ↗Grok 4.6
Cached input / 1M$0.1 ↗$0.5 ↗GPT-6.1 Sol
Input + output / 1M
Lower is cheaper. List prices; batch, tool and regional fees excluded.
$12$8Grok 4.6
Modalitiestext · visiontext · visionTie
Released2026-09-292026-08-12—
Cited benchmark scores2443—
Reliability

Provider status

All providers →
More matchups

Grok 4.6 vs …

Built by Respan
Which one wins on your data?

Public benchmarks are a starting point. Run GPT-6.1 Sol and Grok 4.6 on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.