ModelsCompareBest forBenchmarksStatusPricingAPI
Benchmarks

Every cited score, ranked by exact variant

Pick a benchmark to see every model with a directly comparable, cited score. Variants that change the setup (tools vs no tools, Pro vs Verified) are kept apart on purpose.

Cited scores
1,328
newest measured 2026-09-17
Benchmark variants
326
kept apart, never blended
Models covered
148 / 164
90% of the catalog
Explainers
10
plain-English, below
Pick a benchmark
319 of 319 variants

GPQA Diamond

Graduate-level, 'Google-proof' multiple-choice questions in biology, physics, and chemistry written by domain experts to be hard to answer even with web search.

1
3
Gemini 3.7 FlashGoogleEpoch run94.8% ↗
4
5
7
8
Gemini 3.6 FlashGoogleEpoch run94.1% ↗
9
Grok 4.6xAIEpoch run94% ↗
10
Claude Opus 5AnthropicEpoch run93.9% ↗
11
12
Kimi K3Moonshot AI93.5% ↗
13
Grok 4.5xAIEpoch run93.4% ↗
14
15
16
GPT-5.4OpenAI92.8% ↗
17
Gemini 3.5 FlashGoogleEpoch run92.8% ↗
18
20
GPT-5.2OpenAI92.4% ↗
21
Dola Seed 2.0 ProByteDance Seed92.4% ↗
22
23
GLM-5.2Z.aiEpoch run91.9% ↗
25
Only scores from this exact variant are ranked together. Rows tagged Epoch run or MathArena are independent evaluations, used when a lab has not published its own; ⚠ means MathArena flagged the model as released after the competition. ↗ opens the source.
Explainers

What each benchmark measures

SWE-bench48 models

Measures real-world software engineering ability by asking a model to resolve actual GitHub issues from open-source repositories, verified against the projects' own tests.

#1Claude Opus 4.888.6%
GPQA111 models

Graduate-level, 'Google-proof' multiple-choice questions in biology, physics, and chemistry written by domain experts to be hard to answer even with web search.

#1GPT-6 Astra96%
MMLU37 models

Massive Multitask Language Understanding: multiple-choice questions across 57 subjects (STEM, humanities, law, and more) testing broad knowledge and reasoning.

#1DeepSeek R190.8%
HumanEval2 models

Measures functional code generation: the model writes Python functions from docstrings, scored by whether the code passes hidden unit tests (pass@k).

#1Codestral86.6%
OTIS Mock AIME 2024-202576 models

45 competition-style problems written by students of the Olympiad Training for Individual Study (OTIS) program, with integer answers from 0 to 999. Epoch AI runs every model itself under one setup, so scores are directly comparable.

#1GPT-6 Astra100%
FrontierMath59 models

Original, unpublished mathematics problems created and verified by expert mathematicians: 295 base problems (Tiers 1-3) plus 43 exceptionally hard Tier 4 problems. Answers are checked automatically, and Epoch AI evaluates models on the private set.

#1GPT-6 Astra93.7%
SimpleQA Verified55 models

A cleaned, verified version of the SimpleQA short-form factuality benchmark: fact-seeking questions a model must answer correctly from its own knowledge, without tools.

#1GPT-6 Astra75.6%
AIME40 models

Problems from the American Invitational Mathematics Examination, a competition-level math benchmark used to test multi-step quantitative reasoning.

#1GPT-5.2100%
LMArena (Chatbot Arena)62 models

A human-preference leaderboard: people compare two anonymous models' answers blind and vote, producing an Elo rating from millions of pairwise votes.

#1Claude Opus 4.61497.5
Humanity's Last Exam29 models

A very hard, broad exam of expert-level questions across many fields, designed to remain difficult for frontier models.

#1Claude Fable 5.160.9%