Compare / head-to-head
Claude Opus 4.8vs
Grok 4.6
Grok 4.6 leads 14 of 27 shared benchmarks. Grok 4.6 is 3.8x cheaper per token. Claude Opus 4.8 has the larger context window (1M tokens).
Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.
Shared benchmarks
13 – 14
Grok 4.6 leads
Cheaper per token
Grok 4.6
3.8x cheaper, input + output
Larger context
Claude Opus 4.8
1M tokens
Providers
2 providers
Anthropic · xAI
Claude Opus 4.8
Anthropic · released 2026-05-28
textvision
Quality
Benchmark matrix
27 shared · 13 only Claude Opus 4.8 · 11 only Grok 4.6| Benchmark | Claude Opus 4.8 | Grok 4.6 | Δ | Edge |
|---|---|---|---|---|
| Reported by both · 27 | ||||
| GPQA Diamond | 93.6% ↗ | 94% ↗epoch run | -0.4 pt | Grok 4.6 |
| LMArena Elo | 1452.7 ↗ | 1453.8 ↗ | -1.1 | Grok 4.6 |
| APEX-Agents (Mercor) | 48.9% ↗ | 65.3% ↗ | -16.4 pt | Grok 4.6 |
| Chess Puzzles (Epoch AI run) | 34% ↗epoch run | 40% ↗epoch run | -6 pt | Grok 4.6 |
| DeepSWE v1.1 (Datacurve) | 59% ↗ | 67.5% ↗ | -8.5 pt | Grok 4.6 |
| EBR-bench (Epoch AI run) | 28.6% ↗epoch run | 30.5% ↗epoch run | -1.9 pt | Grok 4.6 |
| FrontierMath Tier 4 v2 (Epoch AI run) | 56.1% ↗epoch run | 31.7% ↗epoch run | +24.4 pt | Claude Opus 4.8 |
| FrontierMath Tiers 1-3 v2 (Epoch AI run) | 80% ↗epoch run | 66% ↗epoch run | +14 pt | Claude Opus 4.8 |
| Furniture Assembly (Epoch AI run) | 42.5% ↗epoch run | 40% ↗epoch run | +2.5 pt | Claude Opus 4.8 |
| LiveBench Agentic Coding (LiveBench) | 50.5% ↗ | 57% ↗ | -6.5 pt | Grok 4.6 |
| LiveBench Coding (LiveBench) | 81.8% ↗ | 76.8% ↗ | +5 pt | Claude Opus 4.8 |
| LiveBench Data Analysis (LiveBench) | 66% ↗ | 73.9% ↗ | -7.9 pt | Grok 4.6 |
| LiveBench Instruction Following (LiveBench) | 72% ↗ | 71.9% ↗ | +0.1 pt | Claude Opus 4.8 |
| LiveBench Language (LiveBench) | 79.7% ↗ | 83.7% ↗ | -4 pt | Grok 4.6 |
| LiveBench Mathematics (LiveBench) | 94.3% ↗ | 92.6% ↗ | +1.7 pt | Claude Opus 4.8 |
| LiveBench Reasoning (LiveBench) | 89.2% ↗ | 90.5% ↗ | -1.3 pt | Grok 4.6 |
| LMArena Agent (LMArena) | 0.0664 ↗ | 0.0128 ↗ | +0.1 | Claude Opus 4.8 |
| LMArena Vision (LMArena) | 1286.4 ↗ | 1263.5 ↗ | +22.9 | Claude Opus 4.8 |
| LMArena WebDev (LMArena) | 1555.5 ↗ | 1619.5 ↗ | -64 | Grok 4.6 |
| Mystery Game Puzzles (Epoch AI run) | 36% ↗epoch run | 34% ↗epoch run | +2 pt | Claude Opus 4.8 |
| OTIS Mock AIME 2024-2025 (Epoch AI run) | 98.3% ↗epoch run | 99.2% ↗epoch run | -0.9 pt | Grok 4.6 |
| SAGE (Vals AI) | 54.8% ↗ | 28.9% ↗ | +25.9 pt | Claude Opus 4.8 |
| SimpleBench (SimpleBench) | 64.8% ↗ | 75.9% ↗ | -11.1 pt | Grok 4.6 |
| SimpleQA Verified | 53% ↗epoch run | 49.3% ↗epoch run | +3.7 pt | Claude Opus 4.8 |
| Terminal-Bench 4.0 (Vals AI) | 23.2% ↗ | 17.2% ↗ | +6 pt | Claude Opus 4.8 |
| Vending-Bench 2 (Andon Labs) | 5787.43 ↗ | 9047.03 ↗ | -3259.6 | Grok 4.6 |
| WeirdML (Håvard Tveit Ihle) | 82.9% ↗ | 67.3% ↗ | +15.6 pt | Claude Opus 4.8 |
| Only Claude Opus 4.8 reports · 13 | ||||
| SWE-bench Verified | 88.6% ↗ | not reported | — | — |
| AIME 2026 | 100% ↗matharena ⚠ | not reported | — | — |
| Humanity's Last Exam (no tools) | 49.8% ↗ | not reported | — | — |
| BrowseComp | 84.3% ↗ | not reported | — | — |
| GSO Opt@1 (GSO) | 47.1% ↗ | not reported | — | — |
| Humanity's Last Exam (with tools) | 57.9% ↗ | not reported | — | — |
| OSWorld-Verified | 83.4% ↗ | not reported | — | — |
| SWE-bench Multilingual | 84.4% ↗ | not reported | — | — |
| SWE-bench Multimodal | 38.4% ↗ | not reported | — | — |
| SWE-Bench Pro | 69.2% ↗ | not reported | — | — |
| tau2-bench Banking Knowledge (Sierra) | 39.7% ↗ | not reported | — | — |
| Terminal-Bench 2.1 | 74.6% ↗ | not reported | — | — |
| Toolathlon-Verified (HKUST) | 76.2% ↗ | not reported | — | — |
| Only Grok 4.6 reports · 11 | ||||
| AA-Briefcase | not reported | 1577 ↗ | — | — |
| APEX-Agents | not reported | 57.5% ↗ | — | — |
| APEX-SWE | not reported | 56.4% ↗ | — | — |
| Artificial Analysis Intelligence Index | not reported | 61 ↗ | — | — |
| CursorBench 3.2 | not reported | 69.9% ↗ | — | — |
| DeepSWE v1.1 | not reported | 65.9% ↗ | — | — |
| FrontierCode v1.1 (Extended) | not reported | 61.3% ↗ | — | — |
| FrontierSWE V2 (Proximal Labs) | not reported | 25.3% ↗ | — | — |
| GDPVal-AA v2 | not reported | 1753 ↗ | — | — |
| Harvey LAB (Vals) | not reported | 15.8% ↗ | — | — |
| Terminal-Bench 3.0 | not reported | 26% ↗ | — | — |
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Claude Opus 4.8 minus Grok 4.6 in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing
Side by side
Official pricing pages and model cards| Spec | Claude Opus 4.8 | Grok 4.6 | Edge |
|---|---|---|---|
| Context window | 1M tokens | 500K tokens | Claude Opus 4.8 |
| Max output | 128K tokens | — | — |
| Input price / 1M | $5 ↗ | $2 ↗ | Grok 4.6 |
| Output price / 1M | $25 ↗ | $6 ↗ | Grok 4.6 |
| Cached input / 1M | $0.5 ↗ | $0.5 ↗ | Tie |
| Input + output / 1M Lower is cheaper. List prices; batch, tool and regional fees excluded. | $30 | $8 | Grok 4.6 |
| Modalities | text · vision | text · vision | Tie |
| Released | 2026-05-28 | 2026-08-12 | — |
| Cited benchmark scores | 47 | 43 | — |
Reliability
Provider status
Live from /statusMore matchups
Claude Opus 4.8 vs …
Models sharing the most benchmarksMore matchups
Grok 4.6 vs …
Models sharing the most benchmarksBuilt by Respan
Which one wins on your data?
Public benchmarks are a starting point. Run Claude Opus 4.8 and Grok 4.6 on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.