ModelsCompareBest forBenchmarksStatusPricingAPI

Kimi K3

Moonshot AI·Operational·Released 2026-07-16textvision
Context1,000k
Max output1,048.576k
Input / 1M$3
Output / 1M$15
Cached in$0.3
Benchmarks59

Benchmarks

Cited public sources · rank is exact-benchmark, apples-to-apples

AA-Briefcase#2/21548
AA-LCR#1/474.7%
Agents' Last Exam#2/228.3%
APEX-Agents#2/341%
AutomationBench#2/230.8%
BabyVision (with Python)85.7%
BrowseComp (1M context, no compaction)90.4%
BrowseComp (context compaction)91.2%
CharXiv Reasoning (no tools)#3/384.8%
CharXiv Reasoning (with Python)91.3%
CorpFin v271.6%
CritPt23.4%
DeepSearchQA (F1)95%
DeepSWE v1.1 (Kimi Code)67.5%
Finance Agent v254.4%
FrontierSWE#1/381.2%
GDPval-AA v2#5/71686
GPQA Diamond#8/7093.5%
Harvey Lab-AA94.6%
JobBench#2/254.3%
Kimi Code Bench 2.072.9%
Legal Research Bench44.2%
MathVision (no tools)94.3%
MathVision (with Python)97.8%
MCP-Atlas#1/384.2%
MCPMark-Verified#1/394.5%
MLS-Bench-Lite#1/348.3%
MMMU-Pro (no tools)#1/681.6%
MMMU-Pro (with Python)83.4%
MMVU#1/282.1%
OfficeQA Pro#1/263.3%
OmniDocBench91.1%
OSWorld 2.058.3%
OSWorld-Verified#1/1284.8%
PerceptionBench58.5%
PostTrainBench#2/336.6%
ProgramBench#1/377.8%
ResearchRubrics#1/376.2%
SaaS-Bench60.1%
SciCode#1/558.7%
SpreadsheetBench 234.8%
SWE-Marathon42%
tau3-Banking33.4%
Terminal-Bench 2.1#2/1788.3%
Toolathlon-Verified76.5%
Video-MME (with subtitles)90%
WorldVQA ForceAnswer51%
ZeroBench pass@5 (no tools)23%
ZeroBench pass@5 (with Python)41%
BrowseComp (1M context, no compaction)90.4%
DeepSWE v1.1#5/1367.3%

Access Kimi K3 and every other model through one endpoint with automatic failover: Respan gateway.