Every cited score, ranked by exact variant
Pick a benchmark to see every model with a directly comparable, cited score. Variants that change the setup (tools vs no tools, Pro vs Verified) are kept apart on purpose.
GPQA Diamond
111 models with a directly comparable cited scoreGraduate-level, 'Google-proof' multiple-choice questions in biology, physics, and chemistry written by domain experts to be hard to answer even with web search.
What each benchmark measures
Measures real-world software engineering ability by asking a model to resolve actual GitHub issues from open-source repositories, verified against the projects' own tests.
Graduate-level, 'Google-proof' multiple-choice questions in biology, physics, and chemistry written by domain experts to be hard to answer even with web search.
Massive Multitask Language Understanding: multiple-choice questions across 57 subjects (STEM, humanities, law, and more) testing broad knowledge and reasoning.
Measures functional code generation: the model writes Python functions from docstrings, scored by whether the code passes hidden unit tests (pass@k).
45 competition-style problems written by students of the Olympiad Training for Individual Study (OTIS) program, with integer answers from 0 to 999. Epoch AI runs every model itself under one setup, so scores are directly comparable.
Original, unpublished mathematics problems created and verified by expert mathematicians: 295 base problems (Tiers 1-3) plus 43 exceptionally hard Tier 4 problems. Answers are checked automatically, and Epoch AI evaluates models on the private set.
A cleaned, verified version of the SimpleQA short-form factuality benchmark: fact-seeking questions a model must answer correctly from its own knowledge, without tools.
Problems from the American Invitational Mathematics Examination, a competition-level math benchmark used to test multi-step quantitative reasoning.
A human-preference leaderboard: people compare two anonymous models' answers blind and vote, producing an Elo rating from millions of pairwise votes.
A very hard, broad exam of expert-level questions across many fields, designed to remain difficult for frontier models.