Compare / head-to-head
Gemini 3.1 Pro Previewvs
Inkling
Gemini 3.1 Pro Preview leads 22 of 24 shared benchmarks. Inkling is 2.8x cheaper per token. Gemini 3.1 Pro Preview has the larger context window (1.0M tokens).
Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.
Shared benchmarks
22 – 1
Gemini 3.1 Pro Preview leads · 1 tied
Cheaper per token
Inkling
2.8x cheaper, input + output
Larger context
Gemini 3.1 Pro Preview
1.0M tokens
Providers
2 providers
Google · Thinking Machines Lab
Gemini 3.1 Pro Preview
Google · released 2026-02-19
textvisionaudio
Quality
Benchmark matrix
24 shared · 12 only Gemini 3.1 Pro Preview · 14 only Inkling| Benchmark | Gemini 3.1 Pro Preview | Inkling | Δ | Edge |
|---|---|---|---|---|
| Reported by both · 24 | ||||
| GPQA Diamond | 94.3% ↗ | 88.3% ↗epoch run | +6 pt | Gemini 3.1 Pro Preview |
| AIME 2026 | 98.3% ↗matharena ⚠ | 97.1% ↗ | +1.2 pt | Gemini 3.1 Pro Preview |
| LMArena Elo | 1480.1 ↗ | 1441.4 ↗ | +38.7 | Gemini 3.1 Pro Preview |
| APEX-Agents (Mercor) | 35.3% ↗ | 33.8% ↗ | +1.5 pt | Gemini 3.1 Pro Preview |
| Chess Puzzles (Epoch AI run) | 55% ↗epoch run | 21% ↗epoch run | +34 pt | Gemini 3.1 Pro Preview |
| FrontierMath Tier 4 v2 (Epoch AI run) | 26.8% ↗epoch run | 4.9% ↗epoch run | +21.9 pt | Gemini 3.1 Pro Preview |
| FrontierMath Tiers 1-3 v2 (Epoch AI run) | 59.6% ↗epoch run | 33.3% ↗epoch run | +26.3 pt | Gemini 3.1 Pro Preview |
| LiveBench Agentic Coding (LiveBench) | 44.1% ↗ | 49.4% ↗ | -5.3 pt | Inkling |
| LiveBench Coding (LiveBench) | 76.5% ↗ | 71% ↗ | +5.5 pt | Gemini 3.1 Pro Preview |
| LiveBench Data Analysis (LiveBench) | 78.5% ↗ | 72.8% ↗ | +5.7 pt | Gemini 3.1 Pro Preview |
| LiveBench Instruction Following (LiveBench) | 79.1% ↗ | 70.1% ↗ | +9 pt | Gemini 3.1 Pro Preview |
| LiveBench Language (LiveBench) | 85.4% ↗ | 73.5% ↗ | +11.9 pt | Gemini 3.1 Pro Preview |
| LiveBench Mathematics (LiveBench) | 91% ↗ | 88.4% ↗ | +2.6 pt | Gemini 3.1 Pro Preview |
| LiveBench Reasoning (LiveBench) | 84% ↗ | 78.3% ↗ | +5.7 pt | Gemini 3.1 Pro Preview |
| LMArena Agent (LMArena) | -0.0771 ↗ | -0.1086 ↗ | 0 | Tie |
| LMArena WebDev (LMArena) | 1446.2 ↗ | 1412.6 ↗ | +33.6 | Gemini 3.1 Pro Preview |
| OTIS Mock AIME 2024-2025 (Epoch AI run) | 95.6% ↗epoch run | 88.9% ↗epoch run | +6.7 pt | Gemini 3.1 Pro Preview |
| SAGE (Vals AI) | 48.7% ↗ | 36.6% ↗ | +12.1 pt | Gemini 3.1 Pro Preview |
| SimpleBench (SimpleBench) | 79.6% ↗ | 50% ↗ | +29.6 pt | Gemini 3.1 Pro Preview |
| SimpleQA Verified | 73.5% ↗epoch run | 43.9% ↗ | +29.6 pt | Gemini 3.1 Pro Preview |
| tau2-bench Banking Knowledge (Sierra) | 26% ↗ | 25% ↗ | +1 pt | Gemini 3.1 Pro Preview |
| Terminal-Bench 4.0 (Vals AI) | 2.5% ↗ | 0.5% ↗ | +2 pt | Gemini 3.1 Pro Preview |
| Toolathlon-Verified (HKUST) | 61.1% ↗ | 45.5% ↗ | +15.6 pt | Gemini 3.1 Pro Preview |
| WeirdML (Håvard Tveit Ihle) | 72.1% ↗ | 32.3% ↗ | +39.8 pt | Gemini 3.1 Pro Preview |
| Only Gemini 3.1 Pro Preview reports · 12 | ||||
| ARC-AGI-2 | 77.1% ↗ | not reported | — | — |
| BALROG (BALROG) | 57% ↗ | not reported | — | — |
| DeepSWE v1.1 (Datacurve) | 11.7% ↗ | not reported | — | — |
| EBR-bench (Epoch AI run) | 14.3% ↗epoch run | not reported | — | — |
| Furniture Assembly (Epoch AI run) | 26.7% ↗epoch run | not reported | — | — |
| GSO Opt@1 (GSO) | 21.6% ↗ | not reported | — | — |
| LMArena Vision (LMArena) | 1279.3 ↗ | not reported | — | — |
| MirrorCode (Epoch AI run) | 8.9% ↗epoch run | not reported | — | — |
| Mystery Game Puzzles (Epoch AI run) | 34% ↗epoch run | not reported | — | — |
| SWE-Bench Pro | 54.2% ↗ | not reported | — | — |
| SWE-bench Verified (Epoch AI run) | 75.6% ↗epoch run | not reported | — | — |
| Vending-Bench 2 (Andon Labs) | 3774.25 ↗ | not reported | — | — |
| Only Inkling reports · 14 | ||||
| SWE-bench Verified | not reported | 77.6% ↗ | — | — |
| Audio MC | not reported | 56.6% ↗ | — | — |
| BrowseComp (w/ ctx management) | not reported | 77.1% ↗ | — | — |
| CharXiv RQ | not reported | 78.1% ↗ | — | — |
| CharXiv RQ (with python) | not reported | 82% ↗ | — | — |
| FrontierSWE V2 (Proximal Labs) | not reported | 4.1% ↗ | — | — |
| Global-MMLU-Lite | not reported | 88.7% ↗ | — | — |
| IFBench | not reported | 79.8% ↗ | — | — |
| MCP Atlas | not reported | 76% ↗ | — | — |
| MMAU | not reported | 77.2% ↗ | — | — |
| SWE-bench Pro (public) | not reported | 54.3% ↗ | — | — |
| Terminal-Bench 2.1 (best harness) | not reported | 63.8% ↗ | — | — |
| Toolathlon Verified | not reported | 45.5% ↗ | — | — |
| VoiceBench | not reported | 91.4% ↗ | — | — |
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Gemini 3.1 Pro Preview minus Inkling in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing
Side by side
Official pricing pages and model cards| Spec | Gemini 3.1 Pro Preview | Inkling | Edge |
|---|---|---|---|
| Context window | 1.0M tokens | 1M tokens | Gemini 3.1 Pro Preview |
| Max output | 66K tokens | — | — |
| Input price / 1M | $2 ↗ | $1 ↗ | Inkling |
| Output price / 1M | $12 ↗ | $4.05 ↗ | Inkling |
| Cached input / 1M | $0.2 ↗ | $0.17 ↗ | Inkling |
| Input + output / 1M Lower is cheaper. List prices; batch, tool and regional fees excluded. | $14 | $5.05 | Inkling |
| Modalities | text · vision · audio | text · vision · audio | Tie |
| Released | 2026-02-19 | 2026-07-15 | — |
| Cited benchmark scores | 43 | 39 | — |
Reliability
Provider status
Live from /statusMore matchups
Gemini 3.1 Pro Preview vs …
Models sharing the most benchmarksMore matchups
Inkling vs …
Models sharing the most benchmarksBuilt by Respan
Which one wins on your data?
Public benchmarks are a starting point. Run Gemini 3.1 Pro Preview and Inkling on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.