Compare / head-to-head
Inklingvs
Kimi K3
Kimi K3 leads 23 of 25 shared benchmarks. Inkling is 3.6x cheaper per token. Kimi K3 has the larger context window (1.0M tokens).
Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.
Shared benchmarks
2 – 23
Kimi K3 leads
Cheaper per token
Inkling
3.6x cheaper, input + output
Larger context
Kimi K3
1.0M tokens
Providers
2 providers
Thinking Machines Lab · Moonshot AI
Inkling
Thinking Machines Lab · released 2026-07-15
textvisionaudio
Quality
Benchmark matrix
25 shared · 13 only Inkling · 55 only Kimi K3| Benchmark | Inkling | Kimi K3 | Δ | Edge |
|---|---|---|---|---|
| Reported by both · 25 | ||||
| GPQA Diamond | 88.3% ↗epoch run | 93.5% ↗ | -5.2 pt | Kimi K3 |
| AIME 2026 | 97.1% ↗ | 96.7% ↗matharena ⚠ | +0.4 pt | Inkling |
| LMArena Elo | 1441.4 ↗ | 1488.4 ↗ | -47 | Kimi K3 |
| APEX-Agents (Mercor) | 33.8% ↗ | 50.6% ↗ | -16.8 pt | Kimi K3 |
| Chess Puzzles (Epoch AI run) | 21% ↗epoch run | 39% ↗epoch run | -18 pt | Kimi K3 |
| FrontierMath Tier 4 v2 (Epoch AI run) | 4.9% ↗epoch run | 39% ↗epoch run | -34.1 pt | Kimi K3 |
| FrontierMath Tiers 1-3 v2 (Epoch AI run) | 33.3% ↗epoch run | 72.2% ↗epoch run | -38.9 pt | Kimi K3 |
| FrontierSWE V2 (Proximal Labs) | 4.1% ↗ | 25.9% ↗ | -21.8 pt | Kimi K3 |
| LiveBench Agentic Coding (LiveBench) | 49.4% ↗ | 62.2% ↗ | -12.8 pt | Kimi K3 |
| LiveBench Coding (LiveBench) | 71% ↗ | 81.4% ↗ | -10.4 pt | Kimi K3 |
| LiveBench Data Analysis (LiveBench) | 72.8% ↗ | 78.7% ↗ | -5.9 pt | Kimi K3 |
| LiveBench Instruction Following (LiveBench) | 70.1% ↗ | 71.4% ↗ | -1.3 pt | Kimi K3 |
| LiveBench Language (LiveBench) | 73.5% ↗ | 85.5% ↗ | -12 pt | Kimi K3 |
| LiveBench Mathematics (LiveBench) | 88.4% ↗ | 84.4% ↗ | +4 pt | Inkling |
| LiveBench Reasoning (LiveBench) | 78.3% ↗ | 90.7% ↗ | -12.4 pt | Kimi K3 |
| LMArena Agent (LMArena) | -0.1086 ↗ | 0.0418 ↗ | -0.2 | Kimi K3 |
| LMArena WebDev (LMArena) | 1412.6 ↗ | 1657.8 ↗ | -245.2 | Kimi K3 |
| OTIS Mock AIME 2024-2025 (Epoch AI run) | 88.9% ↗epoch run | 97.2% ↗epoch run | -8.3 pt | Kimi K3 |
| SAGE (Vals AI) | 36.6% ↗ | 52.8% ↗ | -16.2 pt | Kimi K3 |
| SimpleBench (SimpleBench) | 50% ↗ | 60.7% ↗ | -10.7 pt | Kimi K3 |
| SimpleQA Verified | 43.9% ↗ | 50.6% ↗epoch run | -6.7 pt | Kimi K3 |
| tau2-bench Banking Knowledge (Sierra) | 25% ↗ | 37.1% ↗ | -12.1 pt | Kimi K3 |
| Terminal-Bench 4.0 (Vals AI) | 0.5% ↗ | 17.2% ↗ | -16.7 pt | Kimi K3 |
| Toolathlon-Verified (HKUST) | 45.5% ↗ | 76.5% ↗ | -31 pt | Kimi K3 |
| WeirdML (Håvard Tveit Ihle) | 32.3% ↗ | 82.6% ↗ | -50.3 pt | Kimi K3 |
| Only Inkling reports · 13 | ||||
| SWE-bench Verified | 77.6% ↗ | not reported | — | — |
| Audio MC | 56.6% ↗ | not reported | — | — |
| BrowseComp (w/ ctx management) | 77.1% ↗ | not reported | — | — |
| CharXiv RQ | 78.1% ↗ | not reported | — | — |
| CharXiv RQ (with python) | 82% ↗ | not reported | — | — |
| Global-MMLU-Lite | 88.7% ↗ | not reported | — | — |
| IFBench | 79.8% ↗ | not reported | — | — |
| MCP Atlas | 76% ↗ | not reported | — | — |
| MMAU | 77.2% ↗ | not reported | — | — |
| SWE-bench Pro (public) | 54.3% ↗ | not reported | — | — |
| Terminal-Bench 2.1 (best harness) | 63.8% ↗ | not reported | — | — |
| Toolathlon Verified | 45.5% ↗ | not reported | — | — |
| VoiceBench | 91.4% ↗ | not reported | — | — |
| Only Kimi K3 reports · 55 | ||||
| Humanity's Last Exam (no tools) | not reported | 43.5% ↗ | — | — |
| AA-Briefcase | not reported | 1548 ↗ | — | — |
| AA-LCR | not reported | 74.7% ↗ | — | — |
| Agents' Last Exam | not reported | 28.3% ↗ | — | — |
| APEX-Agents | not reported | 41% ↗ | — | — |
| AutomationBench | not reported | 30.8% ↗ | — | — |
| BabyVision (with Python) | not reported | 85.7% ↗ | — | — |
| BrowseComp (1M context, no compaction) | not reported | 90.4% ↗ | — | — |
| BrowseComp (context compaction) | not reported | 91.2% ↗ | — | — |
| CharXiv Reasoning (no tools) | not reported | 84.8% ↗ | — | — |
| CharXiv Reasoning (with Python) | not reported | 91.3% ↗ | — | — |
| CorpFin v2 | not reported | 71.6% ↗ | — | — |
| CritPt | not reported | 23.4% ↗ | — | — |
| DeepSearchQA (F1) | not reported | 95% ↗ | — | — |
| DeepSWE v1.1 | not reported | 67.3% ↗ | — | — |
| DeepSWE v1.1 (Datacurve) | not reported | 68.5% ↗ | — | — |
| DeepSWE v1.1 (Kimi Code) | not reported | 67.5% ↗ | — | — |
| Finance Agent v2 | not reported | 54.4% ↗ | — | — |
| FrontierSWE | not reported | 81.2% ↗ | — | — |
| Furniture Assembly (Epoch AI run) | not reported | 34.2% ↗epoch run | — | — |
| GDPval-AA v2 | not reported | 1686 ↗ | — | — |
| Harvey Lab-AA | not reported | 94.6% ↗ | — | — |
| Humanity's Last Exam (with tools) | not reported | 56% ↗ | — | — |
| JobBench | not reported | 54.3% ↗ | — | — |
| Kimi Code Bench 2.0 | not reported | 72.9% ↗ | — | — |
| Legal Research Bench | not reported | 44.2% ↗ | — | — |
| MathVision (no tools) | not reported | 94.3% ↗ | — | — |
| MathVision (with Python) | not reported | 97.8% ↗ | — | — |
| MCP-Atlas | not reported | 84.2% ↗ | — | — |
| MCPMark-Verified | not reported | 94.5% ↗ | — | — |
| MLS-Bench-Lite | not reported | 48.3% ↗ | — | — |
| MMMU-Pro (no tools) | not reported | 81.6% ↗ | — | — |
| MMMU-Pro (with Python) | not reported | 83.4% ↗ | — | — |
| MMVU | not reported | 82.1% ↗ | — | — |
| Mystery Game Puzzles (Epoch AI run) | not reported | 26% ↗epoch run | — | — |
| OfficeQA Pro | not reported | 63.3% ↗ | — | — |
| OmniDocBench | not reported | 91.1% ↗ | — | — |
| OSWorld 2.0 | not reported | 58.3% ↗ | — | — |
| OSWorld-Verified | not reported | 84.8% ↗ | — | — |
| PerceptionBench | not reported | 58.5% ↗ | — | — |
| PostTrainBench | not reported | 36.6% ↗ | — | — |
| ProgramBench | not reported | 77.8% ↗ | — | — |
| ResearchRubrics | not reported | 76.2% ↗ | — | — |
| SaaS-Bench | not reported | 60.1% ↗ | — | — |
| SciCode | not reported | 58.7% ↗ | — | — |
| SpreadsheetBench 2 | not reported | 34.8% ↗ | — | — |
| SWE-Marathon | not reported | 42% ↗ | — | — |
| tau3-Banking | not reported | 33.4% ↗ | — | — |
| Terminal-Bench 2.1 | not reported | 88.3% ↗ | — | — |
| Toolathlon-Verified | not reported | 76.5% ↗ | — | — |
| Vending-Bench 2 (Andon Labs) | not reported | 5165.04 ↗ | — | — |
| Video-MME (with subtitles) | not reported | 90% ↗ | — | — |
| WorldVQA ForceAnswer | not reported | 51% ↗ | — | — |
| ZeroBench pass@5 (no tools) | not reported | 23% ↗ | — | — |
| ZeroBench pass@5 (with Python) | not reported | 41% ↗ | — | — |
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Inkling minus Kimi K3 in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing
Side by side
Official pricing pages and model cards| Spec | Inkling | Kimi K3 | Edge |
|---|---|---|---|
| Context window | 1M tokens | 1.0M tokens | Kimi K3 |
| Max output | — | 1.0M tokens | — |
| Input price / 1M | $1 ↗ | $3 ↗ | Inkling |
| Output price / 1M | $4.05 ↗ | $15 ↗ | Inkling |
| Cached input / 1M | $0.17 ↗ | $0.3 ↗ | Inkling |
| Input + output / 1M Lower is cheaper. List prices; batch, tool and regional fees excluded. | $5.05 | $18 | Inkling |
| Modalities | text · vision · audio | text · vision | Inkling |
| Released | 2026-07-15 | 2026-07-16 | — |
| Cited benchmark scores | 39 | 87 | — |
Reliability
Provider status
Live from /statusMore matchups
Inkling vs …
Models sharing the most benchmarksMore matchups
Kimi K3 vs …
Models sharing the most benchmarksBuilt by Respan
Which one wins on your data?
Public benchmarks are a starting point. Run Inkling and Kimi K3 on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.