Compare / head-to-head

Gemini 3.1 Pro PreviewvsKimi K2.5

Gemini 3.1 Pro Preview leads 10 of 11 shared benchmarks. Gemini 3.1 Pro Preview has the larger context window (1.0M tokens).

Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.

Shared benchmarks
10 – 1
Gemini 3.1 Pro Preview leads
Cheaper per token
—
list price, input + output
Larger context
Gemini 3.1 Pro Preview
1.0M tokens
Providers
2 providers
Google · Moonshot AI
Gemini 3.1 Pro Preview
Google · released 2026-02-19
textvisionaudio
Context
1.0M
Max out
66K
Input /1M
$2 ↗
Output /1M
$12 ↗
Cached /1M
$0.2
Scores
43 · 12 core
Kimi K2.5
Moonshot AI · released 2026-01-27
textvisionvideo
Context
256K
Max out
—
Input /1M
—
Output /1M
—
Cached /1M
—
Scores
26 · 14 core
Quality

Benchmark matrix

BenchmarkGemini 3.1 Pro PreviewKimi K2.5ΔEdge
Reported by both · 11
LMArena Elo1480.1 ↗1450.1 ↗+30Gemini 3.1 Pro Preview
LMArena Vision (LMArena)1279.3 ↗1252.2 ↗+27.1Gemini 3.1 Pro Preview
LMArena WebDev (LMArena)1446.2 ↗1435.8 ↗+10.4Gemini 3.1 Pro Preview
SAGE (Vals AI)48.7% ↗49.9% ↗-1.2 ptKimi K2.5
SimpleBench (SimpleBench)79.6% ↗46.8% ↗+32.8 ptGemini 3.1 Pro Preview
SimpleQA Verified73.5% ↗epoch run34.3% ↗epoch run+39.2 ptGemini 3.1 Pro Preview
SWE-Bench Pro54.2% ↗50.7% ↗+3.5 ptGemini 3.1 Pro Preview
SWE-bench Verified (Epoch AI run)75.6% ↗epoch run73.8% ↗epoch run+1.8 ptGemini 3.1 Pro Preview
Toolathlon-Verified (HKUST)61.1% ↗33% ↗+28.1 ptGemini 3.1 Pro Preview
Vending-Bench 2 (Andon Labs)3774.25 ↗1198.46 ↗+2575.8Gemini 3.1 Pro Preview
WeirdML (Håvard Tveit Ihle)72.1% ↗45.6% ↗+26.5 ptGemini 3.1 Pro Preview
Only Gemini 3.1 Pro Preview reports · 25
GPQA Diamond94.3% ↗not reported——
AIME 202698.3% ↗matharena ⚠not reported——
APEX-Agents (Mercor)35.3% ↗not reported——
ARC-AGI-277.1% ↗not reported——
BALROG (BALROG)57% ↗not reported——
Chess Puzzles (Epoch AI run)55% ↗epoch runnot reported——
DeepSWE v1.1 (Datacurve)11.7% ↗not reported——
EBR-bench (Epoch AI run)14.3% ↗epoch runnot reported——
FrontierMath Tier 4 v2 (Epoch AI run)26.8% ↗epoch runnot reported——
FrontierMath Tiers 1-3 v2 (Epoch AI run)59.6% ↗epoch runnot reported——
Furniture Assembly (Epoch AI run)26.7% ↗epoch runnot reported——
GSO Opt@1 (GSO)21.6% ↗not reported——
LiveBench Agentic Coding (LiveBench)44.1% ↗not reported——
LiveBench Coding (LiveBench)76.5% ↗not reported——
LiveBench Data Analysis (LiveBench)78.5% ↗not reported——
LiveBench Instruction Following (LiveBench)79.1% ↗not reported——
LiveBench Language (LiveBench)85.4% ↗not reported——
LiveBench Mathematics (LiveBench)91% ↗not reported——
LiveBench Reasoning (LiveBench)84% ↗not reported——
LMArena Agent (LMArena)-0.0771 ↗not reported——
MirrorCode (Epoch AI run)8.9% ↗epoch runnot reported——
Mystery Game Puzzles (Epoch AI run)34% ↗epoch runnot reported——
OTIS Mock AIME 2024-2025 (Epoch AI run)95.6% ↗epoch runnot reported——
tau2-bench Banking Knowledge (Sierra)26% ↗not reported——
Terminal-Bench 4.0 (Vals AI)2.5% ↗not reported——
Only Kimi K2.5 reports · 15
SWE-bench Verifiednot reported76.8% ↗——
MMLU-Pronot reported87.1% ↗——
AIME 2025not reported96.1% ↗——
BrowseCompnot reported60.6% ↗——
GPQA-Diamondnot reported87.6% ↗——
HLE-Fullnot reported30.1% ↗——
HLE-Full (w/ tools)not reported50.2% ↗——
HMMT 2025 (Feb)not reported95.4% ↗——
IMO-AnswerBenchnot reported81.8% ↗——
LiveCodeBench (v6)not reported85% ↗——
MMMU-Pronot reported78.5% ↗——
OSWorld-Verified (XLANG)not reported63.3% ↗——
SWE-Bench Multilingualnot reported73% ↗——
Terminal Bench 2.0not reported50.8% ↗——
VideoMMMUnot reported86.6% ↗——
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Gemini 3.1 Pro Preview minus Kimi K2.5 in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing

Side by side

SpecGemini 3.1 Pro PreviewKimi K2.5Edge
Context window1.0M tokens256K tokensGemini 3.1 Pro Preview
Max output66K tokens——
Input price / 1M$2 ↗——
Output price / 1M$12 ↗——
Cached input / 1M$0.2 ↗——
Input + output / 1M
Lower is cheaper. List prices; batch, tool and regional fees excluded.
$14——
Modalitiestext · vision · audiotext · vision · videoTie
Released2026-02-192026-01-27—
Cited benchmark scores4326—
Reliability

Provider status

All providers →
More matchups

Gemini 3.1 Pro Preview vs …

Built by Respan
Which one wins on your data?

Public benchmarks are a starting point. Run Gemini 3.1 Pro Preview and Kimi K2.5 on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.