Claude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/MClaude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/M

SWE-bench Leaderboard

Real GitHub issues solved end-to-end. · Prices checked 12 August 2026

SWE-bench Verified gives a model a real repository and a real GitHub issue, and asks for a patch that passes the project’s existing tests. It is the hardest of the benchmarks we track and the closest to what a coding agent is actually asked to do.

Scores are much lower here than on code-generation benchmarks like HumanEval, and that gap is the point: writing a correct function is a different task from navigating an unfamiliar codebase to fix a bug.

Leaderboard

29 of 29 scored

Claude Fable 5 leads on SWE-bench Verified at 91%. The median across the 29 models we hold a score for is 49%. The cheapest model on this board is Gemini 2.5 Flash at $0.285/M blended, scoring 48%.

#ModelProviderScoreInput /1MOutput /1MBlended 70/30Context
1Claude Fable 5Anthropic91%$10.00$50.00$22.001,000,000
2Claude Opus 4.8Anthropic89%$5.00$25.00$11.001,000,000
3GPT-5OpenAI75%$1.25$10.00$3.88400,000
4Claude Opus 4.6Anthropic72%$5.00$25.00$11.001,000,000
5o3OpenAI71%$10.00$40.00$19.00200,000
6Claude Opus 4.5Anthropic68%$15.00$75.00$33.00200,000
7Claude Sonnet 4.6Anthropic65%$3.00$15.00$6.601,000,000
8Gemini 2.5 ProGoogle64%$1.25$10.00$3.881,000,000
9Claude Sonnet 4.5Anthropic62%$3.00$15.00$6.60200,000
10o4-miniOpenAI60%$1.10$4.40$2.09200,000
11GPT-4.1OpenAI55%$2.00$8.00$3.801,000,000
12DeepSeek-R1-0528DeepSeek51%$0.550$2.19$1.0464,000
13Claude Sonnet 3.7Anthropic49%$3.00$15.00$6.60200,000
14Claude Sonnet 3.5Anthropic49%$3.00$15.00$6.60200,000
15o3-miniOpenAI49%$1.10$4.40$2.09200,000
16DeepSeek-R1DeepSeek49%$0.550$2.19$1.0464,000
17DeepSeek-R1 (Groq)Groq49%$0.750$0.990$0.822128,000
18DeepSeek-R1 (Together)Together AI49%$3.00$7.00$4.2064,000
19DeepSeek-R1 (fw)Fireworks AI49%$3.00$8.00$4.5064,000
20DeepSeek-R1 (DI)DeepInfra49%$0.550$2.19$1.0464,000
21DeepSeek-R1 (Cerebras)Cerebras49%$0.550$0.990$0.68264,000
22Gemini 2.5 FlashGoogle48%$0.150$0.600$0.2851,000,000
23Grok 3xAI45%$3.00$15.00$6.60131,000
24DeepSeek-V3-0324DeepSeek44%$0.270$1.10$0.519128,000
25DeepSeek-V3DeepSeek42%$0.270$1.10$0.51964,000
26GPT-4.1 miniOpenAI40%$0.400$1.60$0.7601,000,000
27Claude Haiku 4.5Anthropic38%$0.800$4.00$1.76200,000
28Llama 4 MaverickMeta36%$0.200$0.600$0.320128,000
29GPT-4oOpenAI33%$2.50$10.00$4.75128,000

About these scores

  • Real GitHub issues solved end-to-end.
  • Scores are as published by the model’s provider or a public leaderboard. Provider-reported figures are self-reported and are not independently re-run by us.
  • Where a hosted variant of an open-weights model has no score of its own, it inherits the base model’s — the weights are the same.
  • Prices beside each score are the provider’s published token rates. How the blended rate is calculated.

Frequently asked questions

What is SWE-bench Verified?

A benchmark built from real GitHub issues in open-source Python projects. The model is given the repository and the issue text, and must produce a patch that makes the project’s existing tests pass. "Verified" refers to the human-validated subset, where each task has been checked to be solvable and correctly specified.

Is a higher SWE-bench score always better?

It is the best public signal for autonomous coding work, but it measures one language and one workflow. A model that scores slightly lower may still be the better choice if it is faster, cheaper, has a larger context window for your codebase, or follows your team’s conventions more reliably, none of which the benchmark captures.

Other leaderboards

Rankings that weigh price too

Leaderboards are recomputed on every deploy from the catalogue. Prices checked 12 August 2026.