Compare / head-to-head
Grok 4.5vs
Kimi K3
Kimi K3 leads 18 of 25 shared benchmarks. Grok 4.5 is 2.3x cheaper per token. Kimi K3 has the larger context window (1.0M tokens).
Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.
Shared benchmarks
6 – 18
Kimi K3 leads · 1 tied
Cheaper per token
Grok 4.5
2.3x cheaper, input + output
Larger context
Kimi K3
1.0M tokens
Providers
2 providers
xAI · Moonshot AI
Grok 4.5
xAI · released 2026-07-08
textvision
Quality
Benchmark matrix
25 shared · 2 only Grok 4.5 · 55 only Kimi K3| Benchmark | Grok 4.5 | Kimi K3 | Δ | Edge |
|---|---|---|---|---|
| Reported by both · 25 | ||||
| GPQA Diamond | 93.4% ↗epoch run | 93.5% ↗ | -0.1 pt | Kimi K3 |
| LMArena Elo | 1450.1 ↗ | 1488.4 ↗ | -38.3 | Kimi K3 |
| APEX-Agents (Mercor) | 56.2% ↗ | 50.6% ↗ | +5.6 pt | Grok 4.5 |
| Chess Puzzles (Epoch AI run) | 36% ↗epoch run | 39% ↗epoch run | -3 pt | Kimi K3 |
| DeepSWE v1.1 (Datacurve) | 53.8% ↗ | 68.5% ↗ | -14.7 pt | Kimi K3 |
| FrontierMath Tier 4 v2 (Epoch AI run) | 24.4% ↗epoch run | 39% ↗epoch run | -14.6 pt | Kimi K3 |
| FrontierMath Tiers 1-3 v2 (Epoch AI run) | 57.2% ↗epoch run | 72.2% ↗epoch run | -15 pt | Kimi K3 |
| Furniture Assembly (Epoch AI run) | 22.5% ↗epoch run | 34.2% ↗epoch run | -11.7 pt | Kimi K3 |
| LiveBench Agentic Coding (LiveBench) | 56.5% ↗ | 62.2% ↗ | -5.7 pt | Kimi K3 |
| LiveBench Coding (LiveBench) | 68.6% ↗ | 81.4% ↗ | -12.8 pt | Kimi K3 |
| LiveBench Data Analysis (LiveBench) | 73% ↗ | 78.7% ↗ | -5.7 pt | Kimi K3 |
| LiveBench Instruction Following (LiveBench) | 71.5% ↗ | 71.4% ↗ | +0.1 pt | Grok 4.5 |
| LiveBench Language (LiveBench) | 82.8% ↗ | 85.5% ↗ | -2.7 pt | Kimi K3 |
| LiveBench Mathematics (LiveBench) | 90.8% ↗ | 84.4% ↗ | +6.4 pt | Grok 4.5 |
| LiveBench Reasoning (LiveBench) | 87.2% ↗ | 90.7% ↗ | -3.5 pt | Kimi K3 |
| LMArena Agent (LMArena) | 0.0121 ↗ | 0.0418 ↗ | 0 | Tie |
| LMArena WebDev (LMArena) | 1551.8 ↗ | 1657.8 ↗ | -106 | Kimi K3 |
| OTIS Mock AIME 2024-2025 (Epoch AI run) | 97.8% ↗epoch run | 97.2% ↗epoch run | +0.6 pt | Grok 4.5 |
| SAGE (Vals AI) | 35% ↗ | 52.8% ↗ | -17.8 pt | Kimi K3 |
| SimpleBench (SimpleBench) | 70% ↗ | 60.7% ↗ | +9.3 pt | Grok 4.5 |
| SimpleQA Verified | 48.3% ↗epoch run | 50.6% ↗epoch run | -2.3 pt | Kimi K3 |
| tau2-bench Banking Knowledge (Sierra) | 47.9% ↗ | 37.1% ↗ | +10.8 pt | Grok 4.5 |
| Terminal-Bench 4.0 (Vals AI) | 8.6% ↗ | 17.2% ↗ | -8.6 pt | Kimi K3 |
| Vending-Bench 2 (Andon Labs) | 3887.43 ↗ | 5165.04 ↗ | -1277.6 | Kimi K3 |
| WeirdML (Håvard Tveit Ihle) | 46.4% ↗ | 82.6% ↗ | -36.2 pt | Kimi K3 |
| Only Grok 4.5 reports · 2 | ||||
| LMArena Vision (LMArena) | 1279 ↗ | not reported | — | — |
| SWE-rebench 2026-05-15 to 2026-07-01 (Nebius) | 63.8% ↗ | not reported | — | — |
| Only Kimi K3 reports · 55 | ||||
| AIME 2026 | not reported | 96.7% ↗matharena ⚠ | — | — |
| Humanity's Last Exam (no tools) | not reported | 43.5% ↗ | — | — |
| AA-Briefcase | not reported | 1548 ↗ | — | — |
| AA-LCR | not reported | 74.7% ↗ | — | — |
| Agents' Last Exam | not reported | 28.3% ↗ | — | — |
| APEX-Agents | not reported | 41% ↗ | — | — |
| AutomationBench | not reported | 30.8% ↗ | — | — |
| BabyVision (with Python) | not reported | 85.7% ↗ | — | — |
| BrowseComp (1M context, no compaction) | not reported | 90.4% ↗ | — | — |
| BrowseComp (context compaction) | not reported | 91.2% ↗ | — | — |
| CharXiv Reasoning (no tools) | not reported | 84.8% ↗ | — | — |
| CharXiv Reasoning (with Python) | not reported | 91.3% ↗ | — | — |
| CorpFin v2 | not reported | 71.6% ↗ | — | — |
| CritPt | not reported | 23.4% ↗ | — | — |
| DeepSearchQA (F1) | not reported | 95% ↗ | — | — |
| DeepSWE v1.1 | not reported | 67.3% ↗ | — | — |
| DeepSWE v1.1 (Kimi Code) | not reported | 67.5% ↗ | — | — |
| Finance Agent v2 | not reported | 54.4% ↗ | — | — |
| FrontierSWE | not reported | 81.2% ↗ | — | — |
| FrontierSWE V2 (Proximal Labs) | not reported | 25.9% ↗ | — | — |
| GDPval-AA v2 | not reported | 1686 ↗ | — | — |
| Harvey Lab-AA | not reported | 94.6% ↗ | — | — |
| Humanity's Last Exam (with tools) | not reported | 56% ↗ | — | — |
| JobBench | not reported | 54.3% ↗ | — | — |
| Kimi Code Bench 2.0 | not reported | 72.9% ↗ | — | — |
| Legal Research Bench | not reported | 44.2% ↗ | — | — |
| MathVision (no tools) | not reported | 94.3% ↗ | — | — |
| MathVision (with Python) | not reported | 97.8% ↗ | — | — |
| MCP-Atlas | not reported | 84.2% ↗ | — | — |
| MCPMark-Verified | not reported | 94.5% ↗ | — | — |
| MLS-Bench-Lite | not reported | 48.3% ↗ | — | — |
| MMMU-Pro (no tools) | not reported | 81.6% ↗ | — | — |
| MMMU-Pro (with Python) | not reported | 83.4% ↗ | — | — |
| MMVU | not reported | 82.1% ↗ | — | — |
| Mystery Game Puzzles (Epoch AI run) | not reported | 26% ↗epoch run | — | — |
| OfficeQA Pro | not reported | 63.3% ↗ | — | — |
| OmniDocBench | not reported | 91.1% ↗ | — | — |
| OSWorld 2.0 | not reported | 58.3% ↗ | — | — |
| OSWorld-Verified | not reported | 84.8% ↗ | — | — |
| PerceptionBench | not reported | 58.5% ↗ | — | — |
| PostTrainBench | not reported | 36.6% ↗ | — | — |
| ProgramBench | not reported | 77.8% ↗ | — | — |
| ResearchRubrics | not reported | 76.2% ↗ | — | — |
| SaaS-Bench | not reported | 60.1% ↗ | — | — |
| SciCode | not reported | 58.7% ↗ | — | — |
| SpreadsheetBench 2 | not reported | 34.8% ↗ | — | — |
| SWE-Marathon | not reported | 42% ↗ | — | — |
| tau3-Banking | not reported | 33.4% ↗ | — | — |
| Terminal-Bench 2.1 | not reported | 88.3% ↗ | — | — |
| Toolathlon-Verified | not reported | 76.5% ↗ | — | — |
| Toolathlon-Verified (HKUST) | not reported | 76.5% ↗ | — | — |
| Video-MME (with subtitles) | not reported | 90% ↗ | — | — |
| WorldVQA ForceAnswer | not reported | 51% ↗ | — | — |
| ZeroBench pass@5 (no tools) | not reported | 23% ↗ | — | — |
| ZeroBench pass@5 (with Python) | not reported | 41% ↗ | — | — |
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Grok 4.5 minus Kimi K3 in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing
Side by side
Official pricing pages and model cards| Spec | Grok 4.5 | Kimi K3 | Edge |
|---|---|---|---|
| Context window | 500K tokens | 1.0M tokens | Kimi K3 |
| Max output | — | 1.0M tokens | — |
| Input price / 1M | $2 ↗ | $3 ↗ | Grok 4.5 |
| Output price / 1M | $6 ↗ | $15 ↗ | Grok 4.5 |
| Cached input / 1M | $0.3 ↗ | $0.3 ↗ | Tie |
| Input + output / 1M Lower is cheaper. List prices; batch, tool and regional fees excluded. | $8 | $18 | Grok 4.5 |
| Modalities | text · vision | text · vision | Tie |
| Released | 2026-07-08 | 2026-07-16 | — |
| Cited benchmark scores | 33 | 87 | — |
Reliability
Provider status
Live from /statusMore matchups
Grok 4.5 vs …
Models sharing the most benchmarksMore matchups
Kimi K3 vs …
Models sharing the most benchmarksBuilt by Respan
Which one wins on your data?
Public benchmarks are a starting point. Run Grok 4.5 and Kimi K3 on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.