Compare / head-to-head

GPT-4o minivsGPT-5

GPT-5 leads 12 of 12 shared benchmarks. GPT-4o mini is 15x cheaper per token. GPT-5 has the larger context window (400K tokens).

Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.

Shared benchmarks
0 – 12
GPT-5 leads
Cheaper per token
GPT-4o mini
15x cheaper, input + output
Larger context
GPT-5
400K tokens
Providers
OpenAI
same provider
GPT-4o mini
OpenAI · released 2024-07-18
textvision
Context
128K
Max out
16K
Input /1M
$0.15 ↗
Output /1M
$0.6 ↗
Cached /1M
$0.075
Scores
19 · 8 core
GPT-5
OpenAI · released 2025-08-07
textvision
Context
400K
Max out
128K
Input /1M
$1.25 ↗
Output /1M
$10 ↗
Cached /1M
$0.125
Scores
27 · 11 core
Quality

Benchmark matrix

BenchmarkGPT-4o miniGPT-5ΔEdge
Reported by both · 12
GPQA Diamond40.2% ↗86.2% ↗epoch run-46 ptGPT-5
LMArena Elo1317.6 ↗1434.5 ↗-116.9GPT-5
BALROG (BALROG)17.4% ↗32.8% ↗-15.4 ptGPT-5
Chess Puzzles (Epoch AI run)0% ↗epoch run37% ↗epoch run-37 ptGPT-5
FrontierMath Tiers 1-3 v2 (Epoch AI run)0.7% ↗epoch run55.4% ↗epoch run-54.7 ptGPT-5
LMArena Vision (LMArena)1097.1 ↗1210.2 ↗-113.1GPT-5
MATH Level 5 (Epoch AI run)52.6% ↗epoch run98.1% ↗epoch run-45.5 ptGPT-5
Mystery Game Puzzles (Epoch AI run)12% ↗epoch run23% ↗epoch run-11 ptGPT-5
OTIS Mock AIME 2024-2025 (Epoch AI run)6.9% ↗epoch run91.4% ↗epoch run-84.5 ptGPT-5
SimpleBench (SimpleBench)10.7% ↗56.7% ↗-46 ptGPT-5
SimpleQA Verified8.3% ↗epoch run50.1% ↗epoch run-41.8 ptGPT-5
WeirdML (Håvard Tveit Ihle)11.8% ↗60.7% ↗-48.9 ptGPT-5
Only GPT-4o mini reports · 2
MMLU82% ↗not reported——
AIME 20248.6% ↗not reported——
Only GPT-5 reports · 8
SWE-bench Verifiednot reported74.9% ↗——
AIME 2025not reported94.6% ↗——
EBR-bench (Epoch AI run)not reported12.7% ↗epoch run——
FrontierMath Tier 4 v2 (Epoch AI run)not reported22% ↗epoch run——
GSO Opt@1 (GSO)not reported5.9% ↗——
LMArena WebDev (LMArena)not reported1417.1 ↗——
SAGE (Vals AI)not reported43.7% ↗——
SWE-bench Verified (Epoch AI run)not reported73.6% ↗epoch run——
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is GPT-4o mini minus GPT-5 in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing

Side by side

SpecGPT-4o miniGPT-5Edge
Context window128K tokens400K tokensGPT-5
Max output16K tokens128K tokensGPT-5
Input price / 1M$0.15 ↗$1.25 ↗GPT-4o mini
Output price / 1M$0.6 ↗$10 ↗GPT-4o mini
Cached input / 1M$0.075 ↗$0.125 ↗GPT-4o mini
Input + output / 1M
Lower is cheaper. List prices; batch, tool and regional fees excluded.
$0.75$11.25GPT-4o mini
Modalitiestext · visiontext · visionTie
Released2024-07-182025-08-07—
Cited benchmark scores1927—
Reliability

OpenAI status

All providers →
Built by Respan
Which one wins on your data?

Public benchmarks are a starting point. Run GPT-4o mini and GPT-5 on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.