The numbers behind
the models.

How good they are, how much they cost, and whether they are up right now. Every score cited to a public source; status from official provider feeds. All of it free over JSON.

models
253
providers
26
cited scores
4,110
benchmarks
635
probed live
11
last update
Oct 5, 2026
Leads on reasoning
GPT-6 Astra
96%GPQA Diamond
Wins at coding
Claude Fable 5
95%SWE-bench Verified
Hardest exam
Claude Fable 5.1
60.9%Humanity's Last Exam
Cheapest in the top 10
Gemini 3.8 Flash
$0.75 / $3.75per 1M in / out
Longest context
Llama 4 Scout
10Mtokens
Most reliable API
Anthropic
100%30-day uptime, probed
Leaderboard

Top models

#ModelGPQA ↓SWE-benchHLEMock AIME$ in / outContext
1GPT-6 AstraOpenAI96%——100%$10 / $501.1M
2Claude Sonnet 5.5Anthropic95.6%——100%$2 / $101M
3GPT-6.1 SolOpenAI95.4%——100%$2 / $101.1M
4Gemini 3.8 FlashGoogle95.3%——98.9%$0.75 / $3.751.0M
5Gemini 3.7 FlashGoogle94.8%——97.2%$0.75 / $3.751.0M
6GPT-5.6 SolOpenAI94.6%——100%$4 / $201.1M
7GPT-5.4 ProOpenAI94.4%—42.7%—$30 / $1801.1M
8GPT-6 SolOpenAI94.3%——100%$2 / $101.1M
9Gemini 3.1 Pro PreviewGoogle94.3%——95.6%$2 / $121.0M
10Claude Opus 4.7Anthropic94.2%87.6%46.9%97.8%$5 / $251M
11Gemini 3.6 FlashGoogle94.1%——94.2%$0.75 / $3.751.0M
12Grok 4.6xAI94%——99.2%$2 / $6500K
13Claude Opus 5Anthropic93.9%—56.6%98.9%$5 / $251M
14Claude Opus 4.8Anthropic93.6%88.6%49.8%98.3%$5 / $251M
15Kimi K3Moonshot AI93.5%—43.5%97.2%$3 / $151.0M
16Grok 4.5xAI93.4%——97.8%$2 / $6500K
17GPT-5.2 ProOpenAI93.2%—36.6%—$21 / $168400K
18GPT-5.6 TerraOpenAI92.9%——99.7%$2 / $121.1M
19GPT-5.4OpenAI92.8%—39.8%97.8%$2.5 / $151.1M
20Gemini 3.5 FlashGoogle92.8%——95.6%$1.5 / $91.0M
Every figure links to its source on the model page. GPQA falls back to Epoch AI's independent run when a lab has not published one; Mock AIME is always Epoch's run. Blanks mean no cited score exists.
Live

Provider status

View all →
Official status feeds polled every five minutes; never measured through the Respan gateway.
Pareto

Quality vs cost

Pricing →
0%25%50%75%100%$0.2$0.5$1$2$5$10$20GPT-6.1 Sol: 95.4% at $12/1MClaude Sonnet 5.5: 95.6% at $12/1MGPT-6 Sol: 94.3% at $12/1MGPT-6 Luna: 90.5% at $0.6/1MClaude Opus 5.5: 90.6% at $24/1MGPT-6 Astra: 96% at $60/1MGPT-5.6 Sol: 94.6% at $24/1MGPT-5.6 Terra: 92.9% at $14/1MGPT-5.6 Luna: 92.3% at $1.4/1MClaude Opus 5: 93.9% at $30/1MClaude Sonnet 5: 90.5% at $12/1MClaude Haiku 4.5: 71.2% at $6/1MGemini 3.8 Flash: 95.3% at $4.5/1MGemini 3.1 Pro Preview: 94.3% at $14/1MGemini 3.5 Flash-Lite: 83.3% at $2.8/1MDeepSeek V4.1 Flash: 90.9% at $1.5/1MDeepSeek V4 Pro: 91.7% at $5.28/1MGrok 4.7: 92.7% at $8/1MGrok 4.6: 94% at $8/1MGrok 4.3: 88.8% at $3.75/1MMistral Small 4: 71.2% at $0.75/1MQwen3.8-Max: 92.6% at $8/1MQwen3.8 27B: 89.2% at $3.5/1MGPT-5.5: 90.7% at $35/1MGPT-5.4: 92.8% at $17.5/1MGPT-5.4 Pro: 94.4% at $210/1MGPT-5.4 mini: 88% at $5.25/1MGPT-5.4 nano: 82.8% at $1.45/1MGPT-5.2: 92.4% at $15.75/1MGPT-5.2 Pro: 93.2% at $189/1MClaude Opus 4.8: 93.6% at $30/1MClaude Opus 4.7: 94.2% at $30/1MClaude Opus 4.6: 91.3% at $30/1MClaude Sonnet 4.6: 89.9% at $18/1MClaude Sonnet 4.5: 83.4% at $18/1MGPT-5.1: 87.6% at $11.25/1MGPT-5: 86.2% at $11.25/1MGPT-5 mini: 81.6% at $2.25/1MGPT-5 nano: 69.4% at $0.45/1MGPT-5 Pro: 88.4% at $135/1Mo3: 81.8% at $10/1MGPT-4.1: 66.3% at $10/1MGPT-4.1 mini: 65% at $2/1MGPT-4.1 nano: 50.3% at $0.5/1MGPT-4o: 46% at $12.5/1MGPT-4o mini: 40.2% at $0.75/1MClaude Opus 4.1: 77.3% at $90/1MClaude Opus 4: 76.3% at $90/1MClaude Sonnet 4: 79.2% at $18/1MClaude 3.7 Sonnet: 79.7% at $18/1MClaude 3.5 Sonnet: 55.3% at $18/1MGemini 2.5 Pro: 86.4% at $11.25/1MGemini 2.5 Flash: 82.8% at $2.8/1MGemini 2.5 Flash-Lite: 66.7% at $0.5/1MMinistral 3 14B: 71.2% at $0.4/1MQwen3-Max: 72.6% at $7.2/1MGrok 4.5: 93.4% at $8/1MHy3: 90.4% at $0.66/1MDola Seed 2.0 Pro: 92.4% at $3.5/1MMiniMax M3: 90.9% at $1.5/1MMinistral 3 8B: 66.8% at $0.3/1MMinistral 3 3B: 53.4% at $0.2/1MGrok 4 Fast Reasoning: 85.7% at $0.7/1MGrok 4.20 Reasoning: 89.3% at $3.75/1MKimi K3: 93.5% at $18/1MKimi K2.7 Code: 87.9% at $4.95/1MKimi K2.6: 90.8% at $4.95/1MGLM-5.3: 90.9% at $5.800000000000001/1MGLM-5.3-Flash: 90.2% at $0.65/1MGLM-5.2: 91.9% at $5.800000000000001/1MGLM-5: 87.8% at $4.2/1MGemini 3.7 Flash: 94.8% at $4.5/1MGemini 3.6 Flash: 94.1% at $4.5/1MGemini 3.5 Flash: 92.8% at $10.5/1MGemini 3.1 Flash-Lite: 86.9% at $1.75/1MQwen3.8 2.4T-A95B: 92.6% at $8/1MQwen3.7-Max: 90.9% at $10/1MQwen3.7-Plus: 87.9% at $2/1MQwen3.7-Flash: 82.3% at $0.16/1MQwen3.6-Plus: 88.4% at $3.5/1MQwen3.6-Flash: 83.3% at $1.75/1MQwen3.5-Plus: 84.8% at $2.8/1MQwen3.5-Flash: 82.3% at $0.5/1MQwen3.5 397B-A17B: 88.4% at $4.2/1MQwen3.5 35B-A3B: 83.5% at $2.25/1MQwen3.5 27B: 85.5% at $2.6999999999999997/1MGLM-5.1: 89.9% at $5.800000000000001/1MGLM-4.7: 83.3% at $2.8000000000000003/1MClaude Fable 5: 85.9% at $60/1MClaude Opus 4.5: 87% at $30/1MGemini 3 Pro Preview: 91.9% at $14/1MInkling: 88.3% at $5.05/1MInkling-Small: 88.5% at $1.5/1MQwen3.6-Max-Preview: 87.4% at $9.1/1MQwen3 235B-A22B Thinking 2507: 80.1% at $2.53/1MQwen3 30B-A3B Thinking 2507: 70.1% at $2.6/1MQwen3 30B-A3B Instruct 2507: 55.6% at $1/1MHy4 preview: 92.3% at $3.335/1MDola Seed 2.0 Lite: 85.1% at $2.25/1MDola Seed 2.0 Mini: 79% at $0.5/1MQwen3.7-FlashGPT-6 LunaGPT-5.6 LunaDola Seed 2.0 P…Gemini 3.8 FlashGPT-6.1 SolClaude Sonnet 5…GPT-6 AstraInput + output price, $ per 1M tokens (log)GPQA Diamond (%)
OpenAIAnthropicGoogleDeepSeekxAIMistral AIQwenTencent Hunyuan

Each dot is a model. The dashed line is the Pareto frontier: nothing beats these on both score and cost.

Decision guides

Best model for…

All guides →
Price

Cheapest APIs

View all →
1
Qwen3.7-FlashQwen$0.03 / $0.13 ↗
2
Command R7BCohere$0.0375 / $0.15 ↗
3
Ministral 3 3BMistral AI$0.1 / $0.1 ↗
4
GLM-4 32B 0414 128KZ.ai$0.1 / $0.1 ↗
5
Amazon Nova LiteAmazon$0.06 / $0.24 ↗
6
Ministral 3 8BMistral AI$0.15 / $0.15 ↗
7
Seed1.6 FlashByteDance Seed$0.075 / $0.3 ↗
8
Ministral 3 14BMistral AI$0.2 / $0.2 ↗
Developers

Free JSON API

Docs →
curl https://llmmetric.com/v1/models?provider=anthropic
curl https://llmmetric.com/v1/compare?a=gpt-6-astra&b=claude-sonnet-5-5
curl https://llmmetric.com/v1/status/openai
/v1/models/v1/models/{id}/v1/benchmarks/v1/pricing/v1/status/v1/status/{provider}/v1/compare
Built by Respan
Call every model here with one API key

Respan is a gateway plus observability and evals. One endpoint, automatic failover across providers, and traces for every request.