SWE-bench Leaderboard
Real GitHub issues solved end-to-end. · Prices checked 12 August 2026
SWE-bench Verified gives a model a real repository and a real GitHub issue, and asks for a patch that passes the project’s existing tests. It is the hardest of the benchmarks we track and the closest to what a coding agent is actually asked to do.
Scores are much lower here than on code-generation benchmarks like HumanEval, and that gap is the point: writing a correct function is a different task from navigating an unfamiliar codebase to fix a bug.
Leaderboard
29 of 29 scoredClaude Fable 5 leads on SWE-bench Verified at 91%. The median across the 29 models we hold a score for is 49%. The cheapest model on this board is Gemini 2.5 Flash at $0.285/M blended, scoring 48%.
| # | Model | Provider | Score | Input /1M | Output /1M | Blended 70/30 | Context |
|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | Anthropic | 91% | $10.00 | $50.00 | $22.00 | 1,000,000 |
| 2 | Claude Opus 4.8 | Anthropic | 89% | $5.00 | $25.00 | $11.00 | 1,000,000 |
| 3 | GPT-5 | OpenAI | 75% | $1.25 | $10.00 | $3.88 | 400,000 |
| 4 | Claude Opus 4.6 | Anthropic | 72% | $5.00 | $25.00 | $11.00 | 1,000,000 |
| 5 | o3 | OpenAI | 71% | $10.00 | $40.00 | $19.00 | 200,000 |
| 6 | Claude Opus 4.5 | Anthropic | 68% | $15.00 | $75.00 | $33.00 | 200,000 |
| 7 | Claude Sonnet 4.6 | Anthropic | 65% | $3.00 | $15.00 | $6.60 | 1,000,000 |
| 8 | Gemini 2.5 Pro | 64% | $1.25 | $10.00 | $3.88 | 1,000,000 | |
| 9 | Claude Sonnet 4.5 | Anthropic | 62% | $3.00 | $15.00 | $6.60 | 200,000 |
| 10 | o4-mini | OpenAI | 60% | $1.10 | $4.40 | $2.09 | 200,000 |
| 11 | GPT-4.1 | OpenAI | 55% | $2.00 | $8.00 | $3.80 | 1,000,000 |
| 12 | DeepSeek-R1-0528 | DeepSeek | 51% | $0.550 | $2.19 | $1.04 | 64,000 |
| 13 | Claude Sonnet 3.7 | Anthropic | 49% | $3.00 | $15.00 | $6.60 | 200,000 |
| 14 | Claude Sonnet 3.5 | Anthropic | 49% | $3.00 | $15.00 | $6.60 | 200,000 |
| 15 | o3-mini | OpenAI | 49% | $1.10 | $4.40 | $2.09 | 200,000 |
| 16 | DeepSeek-R1 | DeepSeek | 49% | $0.550 | $2.19 | $1.04 | 64,000 |
| 17 | DeepSeek-R1 (Groq) | Groq | 49% | $0.750 | $0.990 | $0.822 | 128,000 |
| 18 | DeepSeek-R1 (Together) | Together AI | 49% | $3.00 | $7.00 | $4.20 | 64,000 |
| 19 | DeepSeek-R1 (fw) | Fireworks AI | 49% | $3.00 | $8.00 | $4.50 | 64,000 |
| 20 | DeepSeek-R1 (DI) | DeepInfra | 49% | $0.550 | $2.19 | $1.04 | 64,000 |
| 21 | DeepSeek-R1 (Cerebras) | Cerebras | 49% | $0.550 | $0.990 | $0.682 | 64,000 |
| 22 | Gemini 2.5 Flash | 48% | $0.150 | $0.600 | $0.285 | 1,000,000 | |
| 23 | Grok 3 | xAI | 45% | $3.00 | $15.00 | $6.60 | 131,000 |
| 24 | DeepSeek-V3-0324 | DeepSeek | 44% | $0.270 | $1.10 | $0.519 | 128,000 |
| 25 | DeepSeek-V3 | DeepSeek | 42% | $0.270 | $1.10 | $0.519 | 64,000 |
| 26 | GPT-4.1 mini | OpenAI | 40% | $0.400 | $1.60 | $0.760 | 1,000,000 |
| 27 | Claude Haiku 4.5 | Anthropic | 38% | $0.800 | $4.00 | $1.76 | 200,000 |
| 28 | Llama 4 Maverick | Meta | 36% | $0.200 | $0.600 | $0.320 | 128,000 |
| 29 | GPT-4o | OpenAI | 33% | $2.50 | $10.00 | $4.75 | 128,000 |
About these scores
- Real GitHub issues solved end-to-end.
- Scores are as published by the model’s provider or a public leaderboard. Provider-reported figures are self-reported and are not independently re-run by us.
- Where a hosted variant of an open-weights model has no score of its own, it inherits the base model’s — the weights are the same.
- Prices beside each score are the provider’s published token rates. How the blended rate is calculated.
Frequently asked questions
What is SWE-bench Verified?
A benchmark built from real GitHub issues in open-source Python projects. The model is given the repository and the issue text, and must produce a patch that makes the project’s existing tests pass. "Verified" refers to the human-validated subset, where each task has been checked to be solvable and correctly specified.
Is a higher SWE-bench score always better?
It is the best public signal for autonomous coding work, but it measures one language and one workflow. A model that scores slightly lower may still be the better choice if it is faster, cheaper, has a larger context window for your codebase, or follows your team’s conventions more reliably, none of which the benchmark captures.
Other leaderboards
Rankings that weigh price too
Leaderboards are recomputed on every deploy from the catalogue. Prices checked 12 August 2026.