Compare / head-to-head

Claude Sonnet 4.5vsGPT-5.3-Codex

GPT-5.3-Codex leads 5 of 5 shared benchmarks. GPT-5.3-Codex is 1.1x cheaper per token. GPT-5.3-Codex has the larger context window (400K tokens).

Benchmarks from cited public sources; pricing from official pages; status from official provider feeds.

Shared benchmarks
0 – 5
GPT-5.3-Codex leads
Cheaper per token
GPT-5.3-Codex
1.1x cheaper, input + output
Larger context
GPT-5.3-Codex
400K tokens
Providers
2 providers
Anthropic · OpenAI
Claude Sonnet 4.5
Anthropic · released 2025-09-29
textvision
Context
200K
Max out
64K
Input /1M
$3 ↗
Output /1M
$15 ↗
Cached /1M
$0.3
Scores
43 · 12 core
GPT-5.3-Codex
OpenAI · released 2026-02-05
textvision
Context
400K
Max out
128K
Input /1M
$1.75 ↗
Output /1M
$14 ↗
Cached /1M
$0.175
Scores
9 · 2 core
Quality

Benchmark matrix

BenchmarkClaude Sonnet 4.5GPT-5.3-CodexΔEdge
Reported by both · 5
OSWorld-Verified61.4% ↗64.7% ↗-3.3 ptGPT-5.3-Codex
SWE-bench Verified (Epoch AI run)71.3% ↗epoch run74.8% ↗epoch run-3.5 ptGPT-5.3-Codex
Terminal-Bench 2.051% ↗77.3% ↗-26.3 ptGPT-5.3-Codex
Vending-Bench 2 (Andon Labs)3838.74 ↗5940.12 ↗-2101.4GPT-5.3-Codex
WeirdML (Håvard Tveit Ihle)47.7% ↗79.3% ↗-31.6 ptGPT-5.3-Codex
Only Claude Sonnet 4.5 reports · 31
SWE-bench Verified77.2% ↗not reported——
GPQA Diamond83.4% ↗not reported——
AIME 202584.2% ↗matharena ⚠not reported——
Humanity's Last Exam (no tools)17.7% ↗not reported——
LMArena Elo1456.6 ↗not reported——
ARC-AGI-2 (Verified)13.6% ↗not reported——
Chess Puzzles (Epoch AI run)12% ↗epoch runnot reported——
EBR-bench (Epoch AI run)2.4% ↗epoch runnot reported——
FrontierMath Tier 4 v2 (Epoch AI run)2.4% ↗epoch runnot reported——
FrontierMath Tiers 1-3 v2 (Epoch AI run)23.9% ↗epoch runnot reported——
GDPval-AA1276 ↗not reported——
GSO Opt@1 (GSO)12.7% ↗not reported——
Humanity's Last Exam (with tools)33.6% ↗not reported——
LMArena WebDev (LMArena)1393 ↗not reported——
MATH Level 5 (Epoch AI run)97.7% ↗epoch runnot reported——
MCP Atlas43.8% ↗not reported——
MMMLU89.5% ↗not reported——
MMMU-Pro (no tools)63.4% ↗not reported——
MMMU-Pro (with tools)68.9% ↗not reported——
Mystery Game Puzzles (Epoch AI run)17% ↗epoch runnot reported——
OSWorld-Verified (XLANG)62.9% ↗not reported——
OTIS Mock AIME 2024-2025 (Epoch AI run)77.8% ↗epoch runnot reported——
SAGE (Vals AI)36.1% ↗not reported——
SimpleBench (SimpleBench)54.3% ↗not reported——
SimpleQA Verified30.7% ↗epoch runnot reported——
tau2-bench Airline (Sierra)72% ↗not reported——
tau2-bench Banking Knowledge (Sierra)25.3% ↗not reported——
Tau2-bench Retail86.2% ↗not reported——
tau2-bench Retail (Sierra)72.4% ↗not reported——
Tau2-bench Telecom98% ↗not reported——
tau2-bench Telecom (Sierra)84.9% ↗not reported——
Only GPT-5.3-Codex reports · 4
Cybersecurity Capture The Flag Challengesnot reported77.6% ↗——
GDPval (wins or ties)not reported70.9% ↗——
SWE-Bench Pro (Public)not reported56.8% ↗——
SWE-Lancer IC Diamondnot reported81.4% ↗——
Scores tagged "epoch" or "matharena" are independent runs, used only where the lab has not published its own; ⚠ marks rows MathArena flags as released after the competition. Higher is better on every row. Δ is Claude Sonnet 4.5 minus GPT-5.3-Codex in the benchmark's own unit. "Not reported" means the lab has not published that figure; it is not a zero. ↗ opens the source.
Specs & pricing

Side by side

SpecClaude Sonnet 4.5GPT-5.3-CodexEdge
Context window200K tokens400K tokensGPT-5.3-Codex
Max output64K tokens128K tokensGPT-5.3-Codex
Input price / 1M$3 ↗$1.75 ↗GPT-5.3-Codex
Output price / 1M$15 ↗$14 ↗GPT-5.3-Codex
Cached input / 1M$0.3 ↗$0.175 ↗GPT-5.3-Codex
Input + output / 1M
Lower is cheaper. List prices; batch, tool and regional fees excluded.
$18$15.75GPT-5.3-Codex
Modalitiestext · visiontext · visionTie
Released2025-09-292026-02-05—
Cited benchmark scores439—
Reliability

Provider status

All providers →
More matchups

Claude Sonnet 4.5 vs …

More matchups

GPT-5.3-Codex vs …

Built by Respan
Which one wins on your data?

Public benchmarks are a starting point. Run Claude Sonnet 4.5 and GPT-5.3-Codex on your own prompts with Respan evals, or route to either through one gateway key with automatic failover.