Compare / head-to-head

Gemini 3.1 Pro PreviewvsInkling

Gemini 3.1 Pro Preview leads 22 of 24 shared benchmarks. Inkling is 2.8x cheaper per token. Gemini 3.1 Pro Preview has the larger context window (1.0M tokens).

Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.

Shared benchmarks
22 – 1
Gemini 3.1 Pro Preview leads · 1 tied
Cheaper per token
Inkling
2.8x cheaper, input + output
Larger context
Gemini 3.1 Pro Preview
1.0M tokens
Providers
2 providers
Google · Thinking Machines Lab
Gemini 3.1 Pro Preview
Google · released 2026-02-19
textvisionaudio
Context
1.0M
Max out
66K
Input /1M
$2 ↗
Output /1M
$12 ↗
Cached /1M
$0.2
Scores
43 · 12 core
Inkling
Thinking Machines Lab · released 2026-07-15
textvisionaudio
Context
1M
Max out
—
Input /1M
$1 ↗
Output /1M
$4.05 ↗
Cached /1M
$0.17
Scores
39 · 12 core
Quality

Benchmark matrix

BenchmarkGemini 3.1 Pro PreviewInklingΔEdge
Reported by both · 24
GPQA Diamond94.3% ↗88.3% ↗epoch run+6 ptGemini 3.1 Pro Preview
AIME 202698.3% ↗matharena ⚠97.1% ↗+1.2 ptGemini 3.1 Pro Preview
LMArena Elo1480.1 ↗1441.4 ↗+38.7Gemini 3.1 Pro Preview
APEX-Agents (Mercor)35.3% ↗33.8% ↗+1.5 ptGemini 3.1 Pro Preview
Chess Puzzles (Epoch AI run)55% ↗epoch run21% ↗epoch run+34 ptGemini 3.1 Pro Preview
FrontierMath Tier 4 v2 (Epoch AI run)26.8% ↗epoch run4.9% ↗epoch run+21.9 ptGemini 3.1 Pro Preview
FrontierMath Tiers 1-3 v2 (Epoch AI run)59.6% ↗epoch run33.3% ↗epoch run+26.3 ptGemini 3.1 Pro Preview
LiveBench Agentic Coding (LiveBench)44.1% ↗49.4% ↗-5.3 ptInkling
LiveBench Coding (LiveBench)76.5% ↗71% ↗+5.5 ptGemini 3.1 Pro Preview
LiveBench Data Analysis (LiveBench)78.5% ↗72.8% ↗+5.7 ptGemini 3.1 Pro Preview
LiveBench Instruction Following (LiveBench)79.1% ↗70.1% ↗+9 ptGemini 3.1 Pro Preview
LiveBench Language (LiveBench)85.4% ↗73.5% ↗+11.9 ptGemini 3.1 Pro Preview
LiveBench Mathematics (LiveBench)91% ↗88.4% ↗+2.6 ptGemini 3.1 Pro Preview
LiveBench Reasoning (LiveBench)84% ↗78.3% ↗+5.7 ptGemini 3.1 Pro Preview
LMArena Agent (LMArena)-0.0771 ↗-0.1086 ↗0Tie
LMArena WebDev (LMArena)1446.2 ↗1412.6 ↗+33.6Gemini 3.1 Pro Preview
OTIS Mock AIME 2024-2025 (Epoch AI run)95.6% ↗epoch run88.9% ↗epoch run+6.7 ptGemini 3.1 Pro Preview
SAGE (Vals AI)48.7% ↗36.6% ↗+12.1 ptGemini 3.1 Pro Preview
SimpleBench (SimpleBench)79.6% ↗50% ↗+29.6 ptGemini 3.1 Pro Preview
SimpleQA Verified73.5% ↗epoch run43.9% ↗+29.6 ptGemini 3.1 Pro Preview
tau2-bench Banking Knowledge (Sierra)26% ↗25% ↗+1 ptGemini 3.1 Pro Preview
Terminal-Bench 4.0 (Vals AI)2.5% ↗0.5% ↗+2 ptGemini 3.1 Pro Preview
Toolathlon-Verified (HKUST)61.1% ↗45.5% ↗+15.6 ptGemini 3.1 Pro Preview
WeirdML (Håvard Tveit Ihle)72.1% ↗32.3% ↗+39.8 ptGemini 3.1 Pro Preview
Only Gemini 3.1 Pro Preview reports · 12
ARC-AGI-277.1% ↗not reported——
BALROG (BALROG)57% ↗not reported——
DeepSWE v1.1 (Datacurve)11.7% ↗not reported——
EBR-bench (Epoch AI run)14.3% ↗epoch runnot reported——
Furniture Assembly (Epoch AI run)26.7% ↗epoch runnot reported——
GSO Opt@1 (GSO)21.6% ↗not reported——
LMArena Vision (LMArena)1279.3 ↗not reported——
MirrorCode (Epoch AI run)8.9% ↗epoch runnot reported——
Mystery Game Puzzles (Epoch AI run)34% ↗epoch runnot reported——
SWE-Bench Pro54.2% ↗not reported——
SWE-bench Verified (Epoch AI run)75.6% ↗epoch runnot reported——
Vending-Bench 2 (Andon Labs)3774.25 ↗not reported——
Only Inkling reports · 14
SWE-bench Verifiednot reported77.6% ↗——
Audio MCnot reported56.6% ↗——
BrowseComp (w/ ctx management)not reported77.1% ↗——
CharXiv RQnot reported78.1% ↗——
CharXiv RQ (with python)not reported82% ↗——
FrontierSWE V2 (Proximal Labs)not reported4.1% ↗——
Global-MMLU-Litenot reported88.7% ↗——
IFBenchnot reported79.8% ↗——
MCP Atlasnot reported76% ↗——
MMAUnot reported77.2% ↗——
SWE-bench Pro (public)not reported54.3% ↗——
Terminal-Bench 2.1 (best harness)not reported63.8% ↗——
Toolathlon Verifiednot reported45.5% ↗——
VoiceBenchnot reported91.4% ↗——
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Gemini 3.1 Pro Preview minus Inkling in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing

Side by side

SpecGemini 3.1 Pro PreviewInklingEdge
Context window1.0M tokens1M tokensGemini 3.1 Pro Preview
Max output66K tokens——
Input price / 1M$2 ↗$1 ↗Inkling
Output price / 1M$12 ↗$4.05 ↗Inkling
Cached input / 1M$0.2 ↗$0.17 ↗Inkling
Input + output / 1M
Lower is cheaper. List prices; batch, tool and regional fees excluded.
$14$5.05Inkling
Modalitiestext · vision · audiotext · vision · audioTie
Released2026-02-192026-07-15—
Cited benchmark scores4339—
Reliability

Provider status

All providers →
More matchups

Gemini 3.1 Pro Preview vs …

More matchups

Inkling vs …

Built by Respan
Which one wins on your data?

Public benchmarks are a starting point. Run Gemini 3.1 Pro Preview and Inkling on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.