HumanEval Leaderboard
Python function synthesis from docstrings. · Prices checked 12 August 2026
HumanEval asks a model to write a self-contained Python function from its docstring, then runs the function against hidden unit tests. It measures code generation in isolation, no repository, no dependencies, no existing tests to satisfy.
Most capable models now score very highly here, which is why SWE-bench Verified has largely replaced it as the benchmark that separates coding models.
Leaderboard
60 of 208 scoredClaude Fable 5 leads on HumanEval at 96%. The median across the 208 models we hold a score for is 76%. The cheapest model on this board is DeepSeek-Coder-V2 at $0.182/M blended, scoring 90%.
| # | Model | Provider | Score | Input /1M | Output /1M | Blended 70/30 | Context |
|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | Anthropic | 96% | $10.00 | $50.00 | $22.00 | 1,000,000 |
| 2 | GPT-5 | OpenAI | 96% | $1.25 | $10.00 | $3.88 | 400,000 |
| 3 | Claude Opus 4.6 | Anthropic | 95% | $5.00 | $25.00 | $11.00 | 1,000,000 |
| 4 | o3 | OpenAI | 95% | $10.00 | $40.00 | $19.00 | 200,000 |
| 5 | Claude Opus 4.5 | Anthropic | 93% | $15.00 | $75.00 | $33.00 | 200,000 |
| 6 | Claude Sonnet 4.6 | Anthropic | 92% | $3.00 | $15.00 | $6.60 | 1,000,000 |
| 7 | Claude Sonnet 3.5 | Anthropic | 92% | $3.00 | $15.00 | $6.60 | 200,000 |
| 8 | o3-mini | OpenAI | 92% | $1.10 | $4.40 | $2.09 | 200,000 |
| 9 | Gemini 2.5 Pro | 92% | $1.25 | $10.00 | $3.88 | 1,000,000 | |
| 10 | Mistral Large 3 | Mistral | 92% | $0.500 | $1.50 | $0.800 | 256,000 |
| 11 | Mistral Large 2 | Mistral | 92% | $2.00 | $6.00 | $3.20 | 128,000 |
| 12 | Nova Premier | Amazon | 92% | $2.50 | $12.50 | $5.50 | 300,000 |
| 13 | Qwen2.5-Coder-32B (Groq) | Groq | 92% | $0.790 | $0.790 | $0.790 | 128,000 |
| 14 | Mistral-Large-2 (NIM) | Nvidia | 92% | $2.00 | $6.00 | $3.20 | 128,000 |
| 15 | Qwen2.5-Coder-32B | Qwen | 92% | $0.700 | $0.700 | $0.700 | 128,000 |
| 16 | Kimi k1.5 | Moonshot AI | 92% | $2.50 | $10.00 | $4.75 | 128,000 |
| 17 | Claude Sonnet 4.5 | Anthropic | 91% | $3.00 | $15.00 | $6.60 | 200,000 |
| 18 | GPT-4.1 | OpenAI | 90% | $2.00 | $8.00 | $3.80 | 1,000,000 |
| 19 | GPT-4o | OpenAI | 90% | $2.50 | $10.00 | $4.75 | 128,000 |
| 20 | o4-mini | OpenAI | 90% | $1.10 | $4.40 | $2.09 | 200,000 |
| 21 | DeepSeek-R1 | DeepSeek | 90% | $0.550 | $2.19 | $1.04 | 64,000 |
| 22 | DeepSeek-R1-0528 | DeepSeek | 90% | $0.550 | $2.19 | $1.04 | 64,000 |
| 23 | DeepSeek-Coder-V2 | DeepSeek | 90% | $0.140 | $0.280 | $0.182 | 128,000 |
| 24 | Pixtral Large | Mistral | 90% | $2.00 | $6.00 | $3.20 | 128,000 |
| 25 | DeepSeek-R1 (Groq) | Groq | 90% | $0.750 | $0.990 | $0.822 | 128,000 |
| 26 | DeepSeek-R1 (Together) | Together AI | 90% | $3.00 | $7.00 | $4.20 | 64,000 |
| 27 | DeepSeek-R1 (fw) | Fireworks AI | 90% | $3.00 | $8.00 | $4.50 | 64,000 |
| 28 | DeepSeek-R1 (DI) | DeepInfra | 90% | $0.550 | $2.19 | $1.04 | 64,000 |
| 29 | DeepSeek-R1 (Cerebras) | Cerebras | 90% | $0.550 | $0.990 | $0.682 | 64,000 |
| 30 | Qwen2.5-Max | Qwen | 90% | $1.60 | $6.40 | $3.04 | 32,000 |
| 31 | o1 | OpenAI | 89% | $15.00 | $60.00 | $28.50 | 200,000 |
| 32 | DeepSeek-V3 | DeepSeek | 89% | $0.270 | $1.10 | $0.519 | 64,000 |
| 33 | DeepSeek-V3-0324 | DeepSeek | 89% | $0.270 | $1.10 | $0.519 | 128,000 |
| 34 | Llama 3.1 405B | Meta | 89% | $0.800 | $0.800 | $0.800 | 128,000 |
| 35 | Nova Pro | Amazon | 89% | $0.800 | $3.20 | $1.52 | 300,000 |
| 36 | Llama 3.1 405B (Groq) | Groq | 89% | $2.99 | $2.99 | $2.99 | 128,000 |
| 37 | Llama 3.1 405B (Together) | Together AI | 89% | $3.50 | $3.50 | $3.50 | 128,000 |
| 38 | Llama 3.1 405B (fw) | Fireworks AI | 89% | $3.00 | $3.00 | $3.00 | 131,000 |
| 39 | Llama 3.1 405B (Rep) | Replicate | 89% | $9.50 | $9.50 | $9.50 | 128,000 |
| 40 | Llama 3.1 405B (DI) | DeepInfra | 89% | $0.800 | $0.800 | $0.800 | 128,000 |
| 41 | Llama-3.1-405B (Cerebras) | Cerebras | 89% | $0.990 | $0.990 | $0.990 | 128,000 |
| 42 | Llama 3.1 405B (Hyp) | Hyperbolic | 89% | $4.00 | $4.00 | $4.00 | 128,000 |
| 43 | Qwen2.5-32B-Instruct | Qwen | 89% | $0.700 | $0.700 | $0.700 | 128,000 |
| 44 | Qwen2.5-Coder-14B | Qwen | 89% | $0.350 | $0.350 | $0.350 | 128,000 |
| 45 | Claude Sonnet 3.7 | Anthropic | 88% | $3.00 | $15.00 | $6.60 | 200,000 |
| 46 | Llama 4 Maverick | Meta | 88% | $0.200 | $0.600 | $0.320 | 128,000 |
| 47 | Sonar Reasoning Pro | Perplexity | 88% | $2.00 | $8.00 | $3.80 | 127,000 |
| 48 | GPT-4o mini | OpenAI | 87% | $0.150 | $0.600 | $0.285 | 128,000 |
| 49 | GPT-4 Turbo | OpenAI | 87% | $10.00 | $30.00 | $16.00 | 128,000 |
| 50 | Gemini 2.5 Flash | 87% | $0.150 | $0.600 | $0.285 | 1,000,000 | |
| 51 | Grok 3 | xAI | 86% | $3.00 | $15.00 | $6.60 | 131,000 |
| 52 | Grok 3 Fast | xAI | 86% | $5.00 | $25.00 | $11.00 | 131,000 |
| 53 | Qwen2.5-72B (Groq) | Groq | 86% | $0.790 | $0.790 | $0.790 | 128,000 |
| 54 | Qwen2.5-72B (Together) | Together AI | 86% | $1.20 | $1.20 | $1.20 | 32,000 |
| 55 | Qwen2.5-72B (fw) | Fireworks AI | 86% | $0.900 | $0.900 | $0.900 | 32,000 |
| 56 | Qwen2.5-72B (DI) | DeepInfra | 86% | $0.350 | $0.400 | $0.365 | 128,000 |
| 57 | Qwen2.5-72B (Hyp) | Hyperbolic | 86% | $0.400 | $0.400 | $0.400 | 128,000 |
| 58 | Llama-3.1-Nemotron-Ultra-253B | Nvidia | 86% | $1.60 | $1.60 | $1.60 | 128,000 |
| 59 | Qwen2.5-72B-Instruct | Qwen | 86% | $1.20 | $1.20 | $1.20 | 128,000 |
| 60 | Codestral | Mistral | 85% | $0.200 | $0.600 | $0.320 | 256,000 |
About these scores
- Python function synthesis from docstrings.
- Scores are as published by the model’s provider or a public leaderboard. Provider-reported figures are self-reported and are not independently re-run by us.
- Where a hosted variant of an open-weights model has no score of its own, it inherits the base model’s — the weights are the same.
- Prices beside each score are the provider’s published token rates. How the blended rate is calculated.
Showing the top 60 of 208 models we hold a HumanEval score for.
Frequently asked questions
What is the difference between HumanEval and SWE-bench?
HumanEval gives a model an isolated function to write from a docstring. SWE-bench Verified gives it a whole repository and a real issue, and requires a patch that passes the project’s existing tests. HumanEval measures code generation; SWE-bench measures software engineering, and scores on it are far lower across the board.
Other leaderboards
Rankings that weigh price too
Leaderboards are recomputed on every deploy from the catalogue. Prices checked 12 August 2026.