ModelsCompareBest forBenchmarksStatusPricingAPI

What is SWE-bench?

Measures real-world software engineering ability by asking a model to resolve actual GitHub issues from open-source repositories, verified against the projects' own tests.

SWE-bench scores by model

1
Claude Opus 4.8Anthropic88.6%
2
Claude Opus 4.7Anthropic87.6%
3
Claude Sonnet 5Anthropic85.2%
4
Claude Opus 4.6Anthropic80.8%
5
GPT-5.2OpenAI80%
6
Claude Sonnet 4.6Anthropic79.6%
7
Hy3Tencent Hunyuan78%
9
GLM-5Z.ai77.8%
10
Mistral Medium 3.5Mistral AI77.6%
11
Claude Sonnet 4.5Anthropic77.2%
12
Qwen3.6 27BQwen77.2%
13
Step 3.7 FlashStepFun76.5%
14
Dola Seed 2.0 ProByteDance Seed76.5%
16
17
GPT-5OpenAI74.9%
18
Claude Opus 4.1Anthropic74.5%
19
Step 3.5 FlashStepFun74.4%
20
MAI-Thinking-1Microsoft73.5%
21
22
Claude Haiku 4.5Anthropic73.3%
23
DeepSeek V3.2DeepSeek73.1%
24
Claude Sonnet 4Anthropic72.7%
25
Claude Opus 4Anthropic72.5%
26
GPT-5 miniOpenAI71%
29
Claude 3.7 SonnetAnthropic70.3%
31
Qwen3-MaxQwen69.6%
33
o3OpenAI69.1%
34
North Mini CodeCohere67.6%
35
gpt-oss-120bOpenAI62.4%
36
gpt-oss-20bOpenAI60.7%
38
LongCat Flash ChatMeituan LongCat60.4%
39
Gemini 2.5 ProGoogle59.6%
40
GPT-5 nanoOpenAI54.7%