Claude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/MClaude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/M

HumanEval Leaderboard

Python function synthesis from docstrings. · Prices checked 12 August 2026

HumanEval asks a model to write a self-contained Python function from its docstring, then runs the function against hidden unit tests. It measures code generation in isolation, no repository, no dependencies, no existing tests to satisfy.

Most capable models now score very highly here, which is why SWE-bench Verified has largely replaced it as the benchmark that separates coding models.

Leaderboard

60 of 208 scored

Claude Fable 5 leads on HumanEval at 96%. The median across the 208 models we hold a score for is 76%. The cheapest model on this board is DeepSeek-Coder-V2 at $0.182/M blended, scoring 90%.

#ModelProviderScoreInput /1MOutput /1MBlended 70/30Context
1Claude Fable 5Anthropic96%$10.00$50.00$22.001,000,000
2GPT-5OpenAI96%$1.25$10.00$3.88400,000
3Claude Opus 4.6Anthropic95%$5.00$25.00$11.001,000,000
4o3OpenAI95%$10.00$40.00$19.00200,000
5Claude Opus 4.5Anthropic93%$15.00$75.00$33.00200,000
6Claude Sonnet 4.6Anthropic92%$3.00$15.00$6.601,000,000
7Claude Sonnet 3.5Anthropic92%$3.00$15.00$6.60200,000
8o3-miniOpenAI92%$1.10$4.40$2.09200,000
9Gemini 2.5 ProGoogle92%$1.25$10.00$3.881,000,000
10Mistral Large 3Mistral92%$0.500$1.50$0.800256,000
11Mistral Large 2Mistral92%$2.00$6.00$3.20128,000
12Nova PremierAmazon92%$2.50$12.50$5.50300,000
13Qwen2.5-Coder-32B (Groq)Groq92%$0.790$0.790$0.790128,000
14Mistral-Large-2 (NIM)Nvidia92%$2.00$6.00$3.20128,000
15Qwen2.5-Coder-32BQwen92%$0.700$0.700$0.700128,000
16Kimi k1.5Moonshot AI92%$2.50$10.00$4.75128,000
17Claude Sonnet 4.5Anthropic91%$3.00$15.00$6.60200,000
18GPT-4.1OpenAI90%$2.00$8.00$3.801,000,000
19GPT-4oOpenAI90%$2.50$10.00$4.75128,000
20o4-miniOpenAI90%$1.10$4.40$2.09200,000
21DeepSeek-R1DeepSeek90%$0.550$2.19$1.0464,000
22DeepSeek-R1-0528DeepSeek90%$0.550$2.19$1.0464,000
23DeepSeek-Coder-V2DeepSeek90%$0.140$0.280$0.182128,000
24Pixtral LargeMistral90%$2.00$6.00$3.20128,000
25DeepSeek-R1 (Groq)Groq90%$0.750$0.990$0.822128,000
26DeepSeek-R1 (Together)Together AI90%$3.00$7.00$4.2064,000
27DeepSeek-R1 (fw)Fireworks AI90%$3.00$8.00$4.5064,000
28DeepSeek-R1 (DI)DeepInfra90%$0.550$2.19$1.0464,000
29DeepSeek-R1 (Cerebras)Cerebras90%$0.550$0.990$0.68264,000
30Qwen2.5-MaxQwen90%$1.60$6.40$3.0432,000
31o1OpenAI89%$15.00$60.00$28.50200,000
32DeepSeek-V3DeepSeek89%$0.270$1.10$0.51964,000
33DeepSeek-V3-0324DeepSeek89%$0.270$1.10$0.519128,000
34Llama 3.1 405BMeta89%$0.800$0.800$0.800128,000
35Nova ProAmazon89%$0.800$3.20$1.52300,000
36Llama 3.1 405B (Groq)Groq89%$2.99$2.99$2.99128,000
37Llama 3.1 405B (Together)Together AI89%$3.50$3.50$3.50128,000
38Llama 3.1 405B (fw)Fireworks AI89%$3.00$3.00$3.00131,000
39Llama 3.1 405B (Rep)Replicate89%$9.50$9.50$9.50128,000
40Llama 3.1 405B (DI)DeepInfra89%$0.800$0.800$0.800128,000
41Llama-3.1-405B (Cerebras)Cerebras89%$0.990$0.990$0.990128,000
42Llama 3.1 405B (Hyp)Hyperbolic89%$4.00$4.00$4.00128,000
43Qwen2.5-32B-InstructQwen89%$0.700$0.700$0.700128,000
44Qwen2.5-Coder-14BQwen89%$0.350$0.350$0.350128,000
45Claude Sonnet 3.7Anthropic88%$3.00$15.00$6.60200,000
46Llama 4 MaverickMeta88%$0.200$0.600$0.320128,000
47Sonar Reasoning ProPerplexity88%$2.00$8.00$3.80127,000
48GPT-4o miniOpenAI87%$0.150$0.600$0.285128,000
49GPT-4 TurboOpenAI87%$10.00$30.00$16.00128,000
50Gemini 2.5 FlashGoogle87%$0.150$0.600$0.2851,000,000
51Grok 3xAI86%$3.00$15.00$6.60131,000
52Grok 3 FastxAI86%$5.00$25.00$11.00131,000
53Qwen2.5-72B (Groq)Groq86%$0.790$0.790$0.790128,000
54Qwen2.5-72B (Together)Together AI86%$1.20$1.20$1.2032,000
55Qwen2.5-72B (fw)Fireworks AI86%$0.900$0.900$0.90032,000
56Qwen2.5-72B (DI)DeepInfra86%$0.350$0.400$0.365128,000
57Qwen2.5-72B (Hyp)Hyperbolic86%$0.400$0.400$0.400128,000
58Llama-3.1-Nemotron-Ultra-253BNvidia86%$1.60$1.60$1.60128,000
59Qwen2.5-72B-InstructQwen86%$1.20$1.20$1.20128,000
60CodestralMistral85%$0.200$0.600$0.320256,000

About these scores

  • Python function synthesis from docstrings.
  • Scores are as published by the model’s provider or a public leaderboard. Provider-reported figures are self-reported and are not independently re-run by us.
  • Where a hosted variant of an open-weights model has no score of its own, it inherits the base model’s — the weights are the same.
  • Prices beside each score are the provider’s published token rates. How the blended rate is calculated.

Showing the top 60 of 208 models we hold a HumanEval score for.

Frequently asked questions

What is the difference between HumanEval and SWE-bench?

HumanEval gives a model an isolated function to write from a docstring. SWE-bench Verified gives it a whole repository and a real issue, and requires a patch that passes the project’s existing tests. HumanEval measures code generation; SWE-bench measures software engineering, and scores on it are far lower across the board.

Other leaderboards

Rankings that weigh price too

Leaderboards are recomputed on every deploy from the catalogue. Prices checked 12 August 2026.