Claude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/MClaude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/M

Best AI Models for Coding

Ranked from provider-published pricing · Prices checked 12 August 2026

SWE-bench Verified measures whether a model can resolve real GitHub issues end to end, reproducing a bug, editing the right files and passing the project’s tests. It is the closest public benchmark to what a coding assistant is actually asked to do.

Prices sit beside each score so capability and cost can be read together. Coding workloads are output-heavy, so the output rate matters more here than the input rate.

The ranking

top 20 of 29

Claude Fable 5 from Anthropic leads this ranking at 91%. The median across the 29 qualifying models is 49%.

#ModelProviderSWE-bench Verified ScoreHumanEvalInput /1MOutput /1MBlended 70/30
1Claude Fable 5Anthropic91%96$10.00$50.00$22.00
2Claude Opus 4.8Anthropic89%$5.00$25.00$11.00
3GPT-5OpenAI75%96$1.25$10.00$3.88
4Claude Opus 4.6Anthropic72%95$5.00$25.00$11.00
5o3OpenAI71%95$10.00$40.00$19.00
6Claude Opus 4.5Anthropic68%93$15.00$75.00$33.00
7Claude Sonnet 4.6Anthropic65%92$3.00$15.00$6.60
8Gemini 2.5 ProGoogle64%92$1.25$10.00$3.88
9Claude Sonnet 4.5Anthropic62%91$3.00$15.00$6.60
10o4-miniOpenAI60%90$1.10$4.40$2.09
11GPT-4.1OpenAI55%90$2.00$8.00$3.80
12DeepSeek-R1-0528DeepSeek51%90$0.550$2.19$1.04
13Claude Sonnet 3.7Anthropic49%88$3.00$15.00$6.60
14Claude Sonnet 3.5Anthropic49%92$3.00$15.00$6.60
15o3-miniOpenAI49%92$1.10$4.40$2.09
16DeepSeek-R1DeepSeek49%90$0.550$2.19$1.04
17DeepSeek-R1 (Groq)Groq49%90$0.750$0.990$0.822
18DeepSeek-R1 (Together)Together AI49%90$3.00$7.00$4.20
19DeepSeek-R1 (fw)Fireworks AI49%90$3.00$8.00$4.50
20DeepSeek-R1 (DI)DeepInfra49%90$0.550$2.19$1.04

How this list is built

  • Ranked by SWE-bench Verified score, highest first. Scores are as published by the model’s provider or a public leaderboard. Where a hosted variant has no score of its own, it inherits the base model’s.
  • Includes models marked available or preview; retired and deprecated SKUs are excluded.
  • Prices are the rates each provider publishes, not estimates. How the blended rate is calculated.

257 otherwise-eligible models were left out because we hold no published value for the figure this list ranks by. They are excluded rather than assumed.

Frequently asked questions

What is the best AI model for coding?

On SWE-bench Verified (resolving real GitHub issues) the models at the top of this table lead. Benchmarks measure a narrow slice of the job, though: latency, context window and how well a model follows your project’s conventions all shape the day-to-day experience and none of them appear in a score.

What is the difference between SWE-bench and HumanEval?

HumanEval asks a model to write a single self-contained Python function from a docstring. SWE-bench Verified gives it an entire repository and a real issue, and requires a patch that passes the project’s existing tests. HumanEval measures code generation; SWE-bench measures software engineering, and scores on it are much lower across the board.

Related rankings

Price it for your workload

Rankings are recomputed on every deploy from the catalogue. Prices checked 12 August 2026.