The numbers behind
the models.
How good they are, how much they cost, and whether they are up right now. Every score cited to a public source; status from official provider feeds. All of it free over JSON.
- models
- 253
- providers
- 26
- cited scores
- 4,110
- benchmarks
- 635
- probed live
- 11
- last update
- Oct 5, 2026
Leads on reasoning
96%GPQA Diamond
Wins at coding
95%SWE-bench Verified
Hardest exam
60.9%Humanity's Last Exam
Cheapest in the top 10
$0.75 / $3.75per 1M in / out
Longest context
10Mtokens
Most reliable API
100%30-day uptime, probed
Leaderboard
Top models
142 models with a cited reasoning figure| # | Model | GPQA ↓ | SWE-bench | HLE | Mock AIME | $ in / out | Context |
|---|---|---|---|---|---|---|---|
| 1 | 96% | — | — | 100% | $10 / $50 | 1.1M | |
| 2 | 95.6% | — | — | 100% | $2 / $10 | 1M | |
| 3 | 95.4% | — | — | 100% | $2 / $10 | 1.1M | |
| 4 | 95.3% | — | — | 98.9% | $0.75 / $3.75 | 1.0M | |
| 5 | 94.8% | — | — | 97.2% | $0.75 / $3.75 | 1.0M | |
| 6 | 94.6% | — | — | 100% | $4 / $20 | 1.1M | |
| 7 | 94.4% | — | 42.7% | — | $30 / $180 | 1.1M | |
| 8 | 94.3% | — | — | 100% | $2 / $10 | 1.1M | |
| 9 | 94.3% | — | — | 95.6% | $2 / $12 | 1.0M | |
| 10 | 94.2% | 87.6% | 46.9% | 97.8% | $5 / $25 | 1M | |
| 11 | 94.1% | — | — | 94.2% | $0.75 / $3.75 | 1.0M | |
| 12 | 94% | — | — | 99.2% | $2 / $6 | 500K | |
| 13 | 93.9% | — | 56.6% | 98.9% | $5 / $25 | 1M | |
| 14 | 93.6% | 88.6% | 49.8% | 98.3% | $5 / $25 | 1M | |
| 15 | 93.5% | — | 43.5% | 97.2% | $3 / $15 | 1.0M | |
| 16 | 93.4% | — | — | 97.8% | $2 / $6 | 500K | |
| 17 | 93.2% | — | 36.6% | — | $21 / $168 | 400K | |
| 18 | 92.9% | — | — | 99.7% | $2 / $12 | 1.1M | |
| 19 | 92.8% | — | 39.8% | 97.8% | $2.5 / $15 | 1.1M | |
| 20 | 92.8% | — | — | 95.6% | $1.5 / $9 | 1.0M |
Every figure links to its source on the model page. GPQA falls back to Epoch AI's independent run when a lab has not published one; Mock AIME is always Epoch's run. Blanks mean no cited score exists.
Live
Provider status
11 probed · last 60 daysOpenAI99.3292%Anthropic100%Google100%DeepSeek100%xAI100%Mistral AI99.3067%Cohere100%Moonshot AI99.8359%MiniMax100%Perplexity100%
Official status feeds polled every five minutes; never measured through the Respan gateway.
Pareto
Quality vs cost
GPQA Diamond vs combined list priceOpenAIAnthropicGoogleDeepSeekxAIMistral AIQwenTencent Hunyuan
Each dot is a model. The dashed line is the Pareto frontier: nothing beats these on both score and cost.
Decision guides
Best model for…
Codingbenchmark
Claude Fable 5
95% SWE-bench Verified
2Claude Opus 4.888.6%
3Claude Opus 4.787.6%
Reasoningbenchmark
GPT-6 Astra
96% GPQA Diamond
2Claude Sonnet 5.595.6%
3GPT-6.1 Sol95.4%
Agents & tool usebenchmark
GPT-5.5
82.7% Terminal-Bench 2.0
2GPT-5.3-Codex77.3%
3GPT-5.475.1%
Long contextcontext
Llama 4 Scout
10M tokens
2Grok 4 Fast Reasoning2M
3Grok 4.1 Fast Reasoning2M
Cheapest APIsprice
Qwen3.7-Flash
$0.03 / $0.13
2Command R7B$0.0375
3Ministral 3 3B$0.1
Visionbenchmark
Dola Seed 2.0 Lite
83.7% MMMU
2Gemini 3.5 Flash83.6%
3Kimi K383.4%
Head-to-head
Popular comparisons
Flagships across the top labsPrice
Cheapest APIs
$ per 1M in / outDevelopers
Free JSON API
No key, CORS on, cached at the edgecurl https://llmmetric.com/v1/models?provider=anthropic curl https://llmmetric.com/v1/compare?a=gpt-6-astra&b=claude-sonnet-5-5 curl https://llmmetric.com/v1/status/openai
/v1/models/v1/models/{id}/v1/benchmarks/v1/pricing/v1/status/v1/status/{provider}/v1/compare
Built by Respan
Call every model here with one API key
Respan is a gateway plus observability and evals. One endpoint, automatic failover across providers, and traces for every request.