ModelsCompareBest forBenchmarksStatusPricingAPI

What is Humanity's Last Exam?

A very hard, broad exam of expert-level questions across many fields, designed to remain difficult for frontier models.

Humanity's Last Exam scores by model

1
Claude Fable 5.1Anthropic60.9%
2
Claude Opus 5Anthropic56.6%
3
Claude Opus 4.8Anthropic49.8%
4
Claude Opus 4.7Anthropic46.9%
5
Kimi K3Moonshot AI43.5%
6
Claude Sonnet 5Anthropic43.2%
7
GPT-5.5 ProOpenAI43.1%
8
GPT-5.4 ProOpenAI42.7%
9
GPT-5.4OpenAI39.8%
10
Hy3Tencent Hunyuan37%
11
GPT-5.2 ProOpenAI36.6%
13
Dola Seed 2.0 ProByteDance Seed33.3%
14
Claude Sonnet 4.6Anthropic33.2%
17
DeepSeek V3.2DeepSeek25.1%
18
Gemini 2.5 ProGoogle21.6%
20
Gemma 4 31BGoogle19.5%
22
Claude Sonnet 4.5Anthropic17.7%
27