Compare / head-to-head

Grok 4.5vsKimi K3

Kimi K3 leads 18 of 25 shared benchmarks. Grok 4.5 is 2.3x cheaper per token. Kimi K3 has the larger context window (1.0M tokens).

Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.

Shared benchmarks
6 – 18
Kimi K3 leads · 1 tied
Cheaper per token
Grok 4.5
2.3x cheaper, input + output
Larger context
Kimi K3
1.0M tokens
Providers
2 providers
xAI · Moonshot AI
Grok 4.5
xAI · released 2026-07-08
textvision
Context
500K
Max out
—
Input /1M
$2 ↗
Output /1M
$6 ↗
Cached /1M
$0.3
Scores
33 · 9 core
Kimi K3
Moonshot AI · released 2026-07-16
textvision
Context
1.0M
Max out
1.0M
Input /1M
$3 ↗
Output /1M
$15 ↗
Cached /1M
$0.3
Scores
87 · 11 core
Quality

Benchmark matrix

BenchmarkGrok 4.5Kimi K3ΔEdge
Reported by both · 25
GPQA Diamond93.4% ↗epoch run93.5% ↗-0.1 ptKimi K3
LMArena Elo1450.1 ↗1488.4 ↗-38.3Kimi K3
APEX-Agents (Mercor)56.2% ↗50.6% ↗+5.6 ptGrok 4.5
Chess Puzzles (Epoch AI run)36% ↗epoch run39% ↗epoch run-3 ptKimi K3
DeepSWE v1.1 (Datacurve)53.8% ↗68.5% ↗-14.7 ptKimi K3
FrontierMath Tier 4 v2 (Epoch AI run)24.4% ↗epoch run39% ↗epoch run-14.6 ptKimi K3
FrontierMath Tiers 1-3 v2 (Epoch AI run)57.2% ↗epoch run72.2% ↗epoch run-15 ptKimi K3
Furniture Assembly (Epoch AI run)22.5% ↗epoch run34.2% ↗epoch run-11.7 ptKimi K3
LiveBench Agentic Coding (LiveBench)56.5% ↗62.2% ↗-5.7 ptKimi K3
LiveBench Coding (LiveBench)68.6% ↗81.4% ↗-12.8 ptKimi K3
LiveBench Data Analysis (LiveBench)73% ↗78.7% ↗-5.7 ptKimi K3
LiveBench Instruction Following (LiveBench)71.5% ↗71.4% ↗+0.1 ptGrok 4.5
LiveBench Language (LiveBench)82.8% ↗85.5% ↗-2.7 ptKimi K3
LiveBench Mathematics (LiveBench)90.8% ↗84.4% ↗+6.4 ptGrok 4.5
LiveBench Reasoning (LiveBench)87.2% ↗90.7% ↗-3.5 ptKimi K3
LMArena Agent (LMArena)0.0121 ↗0.0418 ↗0Tie
LMArena WebDev (LMArena)1551.8 ↗1657.8 ↗-106Kimi K3
OTIS Mock AIME 2024-2025 (Epoch AI run)97.8% ↗epoch run97.2% ↗epoch run+0.6 ptGrok 4.5
SAGE (Vals AI)35% ↗52.8% ↗-17.8 ptKimi K3
SimpleBench (SimpleBench)70% ↗60.7% ↗+9.3 ptGrok 4.5
SimpleQA Verified48.3% ↗epoch run50.6% ↗epoch run-2.3 ptKimi K3
tau2-bench Banking Knowledge (Sierra)47.9% ↗37.1% ↗+10.8 ptGrok 4.5
Terminal-Bench 4.0 (Vals AI)8.6% ↗17.2% ↗-8.6 ptKimi K3
Vending-Bench 2 (Andon Labs)3887.43 ↗5165.04 ↗-1277.6Kimi K3
WeirdML (Håvard Tveit Ihle)46.4% ↗82.6% ↗-36.2 ptKimi K3
Only Grok 4.5 reports · 2
LMArena Vision (LMArena)1279 ↗not reported——
SWE-rebench 2026-05-15 to 2026-07-01 (Nebius)63.8% ↗not reported——
Only Kimi K3 reports · 55
AIME 2026not reported96.7% ↗matharena ⚠——
Humanity's Last Exam (no tools)not reported43.5% ↗——
AA-Briefcasenot reported1548 ↗——
AA-LCRnot reported74.7% ↗——
Agents' Last Examnot reported28.3% ↗——
APEX-Agentsnot reported41% ↗——
AutomationBenchnot reported30.8% ↗——
BabyVision (with Python)not reported85.7% ↗——
BrowseComp (1M context, no compaction)not reported90.4% ↗——
BrowseComp (context compaction)not reported91.2% ↗——
CharXiv Reasoning (no tools)not reported84.8% ↗——
CharXiv Reasoning (with Python)not reported91.3% ↗——
CorpFin v2not reported71.6% ↗——
CritPtnot reported23.4% ↗——
DeepSearchQA (F1)not reported95% ↗——
DeepSWE v1.1not reported67.3% ↗——
DeepSWE v1.1 (Kimi Code)not reported67.5% ↗——
Finance Agent v2not reported54.4% ↗——
FrontierSWEnot reported81.2% ↗——
FrontierSWE V2 (Proximal Labs)not reported25.9% ↗——
GDPval-AA v2not reported1686 ↗——
Harvey Lab-AAnot reported94.6% ↗——
Humanity's Last Exam (with tools)not reported56% ↗——
JobBenchnot reported54.3% ↗——
Kimi Code Bench 2.0not reported72.9% ↗——
Legal Research Benchnot reported44.2% ↗——
MathVision (no tools)not reported94.3% ↗——
MathVision (with Python)not reported97.8% ↗——
MCP-Atlasnot reported84.2% ↗——
MCPMark-Verifiednot reported94.5% ↗——
MLS-Bench-Litenot reported48.3% ↗——
MMMU-Pro (no tools)not reported81.6% ↗——
MMMU-Pro (with Python)not reported83.4% ↗——
MMVUnot reported82.1% ↗——
Mystery Game Puzzles (Epoch AI run)not reported26% ↗epoch run——
OfficeQA Pronot reported63.3% ↗——
OmniDocBenchnot reported91.1% ↗——
OSWorld 2.0not reported58.3% ↗——
OSWorld-Verifiednot reported84.8% ↗——
PerceptionBenchnot reported58.5% ↗——
PostTrainBenchnot reported36.6% ↗——
ProgramBenchnot reported77.8% ↗——
ResearchRubricsnot reported76.2% ↗——
SaaS-Benchnot reported60.1% ↗——
SciCodenot reported58.7% ↗——
SpreadsheetBench 2not reported34.8% ↗——
SWE-Marathonnot reported42% ↗——
tau3-Bankingnot reported33.4% ↗——
Terminal-Bench 2.1not reported88.3% ↗——
Toolathlon-Verifiednot reported76.5% ↗——
Toolathlon-Verified (HKUST)not reported76.5% ↗——
Video-MME (with subtitles)not reported90% ↗——
WorldVQA ForceAnswernot reported51% ↗——
ZeroBench pass@5 (no tools)not reported23% ↗——
ZeroBench pass@5 (with Python)not reported41% ↗——
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Grok 4.5 minus Kimi K3 in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing

Side by side

SpecGrok 4.5Kimi K3Edge
Context window500K tokens1.0M tokensKimi K3
Max output—1.0M tokens—
Input price / 1M$2 ↗$3 ↗Grok 4.5
Output price / 1M$6 ↗$15 ↗Grok 4.5
Cached input / 1M$0.3 ↗$0.3 ↗Tie
Input + output / 1M
Lower is cheaper. List prices; batch, tool and regional fees excluded.
$8$18Grok 4.5
Modalitiestext · visiontext · visionTie
Released2026-07-082026-07-16—
Cited benchmark scores3387—
Reliability

Provider status

All providers →
More matchups

Kimi K3 vs …

Built by Respan
Which one wins on your data?

Public benchmarks are a starting point. Run Grok 4.5 and Kimi K3 on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.