MMLU Leaderboard
Multitask academic knowledge across 57 subjects. · Prices checked 12 August 2026
MMLU tests multiple-choice knowledge across 57 subjects, from elementary mathematics to professional law. It was the standard general-capability benchmark for years and remains a useful breadth check.
It is now close to saturated at the top: the strongest models cluster within a few points of each other, so small differences here mean much less than they did when the benchmark was new.
Leaderboard
60 of 200 scoredGPT-5 leads on MMLU at 91%. The median across the 200 models we hold a score for is 79%. The cheapest model on this board is Llama 3.3 70B at $0.281/M blended, scoring 86%.
| # | Model | Provider | Score | Input /1M | Output /1M | Blended 70/30 | Context |
|---|---|---|---|---|---|---|---|
| 1 | GPT-5 | OpenAI | 91% | $1.25 | $10.00 | $3.88 | 400,000 |
| 2 | o1 | OpenAI | 91% | $15.00 | $60.00 | $28.50 | 200,000 |
| 3 | Claude Fable 5 | Anthropic | 90% | $10.00 | $50.00 | $22.00 | 1,000,000 |
| 4 | o3 | OpenAI | 90% | $10.00 | $40.00 | $19.00 | 200,000 |
| 5 | DeepSeek-R1 | DeepSeek | 90% | $0.550 | $2.19 | $1.04 | 64,000 |
| 6 | DeepSeek-R1-0528 | DeepSeek | 90% | $0.550 | $2.19 | $1.04 | 64,000 |
| 7 | DeepSeek-R1 (Groq) | Groq | 90% | $0.750 | $0.990 | $0.822 | 128,000 |
| 8 | DeepSeek-R1 (Together) | Together AI | 90% | $3.00 | $7.00 | $4.20 | 64,000 |
| 9 | DeepSeek-R1 (fw) | Fireworks AI | 90% | $3.00 | $8.00 | $4.50 | 64,000 |
| 10 | DeepSeek-R1 (DI) | DeepInfra | 90% | $0.550 | $2.19 | $1.04 | 64,000 |
| 11 | DeepSeek-R1 (Cerebras) | Cerebras | 90% | $0.550 | $0.990 | $0.682 | 64,000 |
| 12 | GPT-4.1 | OpenAI | 89% | $2.00 | $8.00 | $3.80 | 1,000,000 |
| 13 | Gemini 2.5 Pro | 89% | $1.25 | $10.00 | $3.88 | 1,000,000 | |
| 14 | Claude Opus 4.6 | Anthropic | 88% | $5.00 | $25.00 | $11.00 | 1,000,000 |
| 15 | GPT-4o | OpenAI | 88% | $2.50 | $10.00 | $4.75 | 128,000 |
| 16 | DeepSeek-V3 | DeepSeek | 88% | $0.270 | $1.10 | $0.519 | 64,000 |
| 17 | DeepSeek-V3-0324 | DeepSeek | 88% | $0.270 | $1.10 | $0.519 | 128,000 |
| 18 | Llama 3.1 405B | Meta | 88% | $0.800 | $0.800 | $0.800 | 128,000 |
| 19 | Nova Premier | Amazon | 88% | $2.50 | $12.50 | $5.50 | 300,000 |
| 20 | Sonar Reasoning Pro | Perplexity | 88% | $2.00 | $8.00 | $3.80 | 127,000 |
| 21 | Llama 3.1 405B (Groq) | Groq | 88% | $2.99 | $2.99 | $2.99 | 128,000 |
| 22 | Llama 3.1 405B (Together) | Together AI | 88% | $3.50 | $3.50 | $3.50 | 128,000 |
| 23 | Llama 3.1 405B (fw) | Fireworks AI | 88% | $3.00 | $3.00 | $3.00 | 131,000 |
| 24 | Llama 3.1 405B (Rep) | Replicate | 88% | $9.50 | $9.50 | $9.50 | 128,000 |
| 25 | Llama 3.1 405B (DI) | DeepInfra | 88% | $0.800 | $0.800 | $0.800 | 128,000 |
| 26 | Llama-3.1-405B (Cerebras) | Cerebras | 88% | $0.990 | $0.990 | $0.990 | 128,000 |
| 27 | Llama 3.1 405B (Hyp) | Hyperbolic | 88% | $4.00 | $4.00 | $4.00 | 128,000 |
| 28 | Llama-3.1-Nemotron-Ultra-253B | Nvidia | 88% | $1.60 | $1.60 | $1.60 | 128,000 |
| 29 | Kimi k1.5 | Moonshot AI | 88% | $2.50 | $10.00 | $4.75 | 128,000 |
| 30 | Claude Opus 4.5 | Anthropic | 87% | $15.00 | $75.00 | $33.00 | 200,000 |
| 31 | Sonar Deep Research | Perplexity | 87% | $2.00 | $8.00 | $3.80 | 127,000 |
| 32 | Qwen2.5-Max | Qwen | 87% | $1.60 | $6.40 | $3.04 | 32,000 |
| 33 | Claude Sonnet 4.6 | Anthropic | 86% | $3.00 | $15.00 | $6.60 | 1,000,000 |
| 34 | GPT-4 Turbo | OpenAI | 86% | $10.00 | $30.00 | $16.00 | 128,000 |
| 35 | o3-mini | OpenAI | 86% | $1.10 | $4.40 | $2.09 | 200,000 |
| 36 | Gemini 1.5 Pro | 86% | $1.25 | $5.00 | $2.38 | 2,000,000 | |
| 37 | Grok 3 | xAI | 86% | $3.00 | $15.00 | $6.60 | 131,000 |
| 38 | Grok 3 Fast | xAI | 86% | $5.00 | $25.00 | $11.00 | 131,000 |
| 39 | Llama 3.3 70B | Meta | 86% | $0.230 | $0.400 | $0.281 | 128,000 |
| 40 | Llama 3.2 90B Vision | Meta | 86% | $0.880 | $0.880 | $0.880 | 128,000 |
| 41 | Nova Pro | Amazon | 86% | $0.800 | $3.20 | $1.52 | 300,000 |
| 42 | Sonar Pro | Perplexity | 86% | $3.00 | $15.00 | $6.60 | 200,000 |
| 43 | Llama 3.3 70B (Groq) | Groq | 86% | $0.590 | $0.790 | $0.650 | 128,000 |
| 44 | Qwen2.5-72B (Groq) | Groq | 86% | $0.790 | $0.790 | $0.790 | 128,000 |
| 45 | Llama 3.3 70B (Together) | Together AI | 86% | $0.880 | $0.880 | $0.880 | 128,000 |
| 46 | Qwen2.5-72B (Together) | Together AI | 86% | $1.20 | $1.20 | $1.20 | 32,000 |
| 47 | Qwen2.5-72B (fw) | Fireworks AI | 86% | $0.900 | $0.900 | $0.900 | 32,000 |
| 48 | Llama 3.3 70B (DI) | DeepInfra | 86% | $0.230 | $0.400 | $0.281 | 128,000 |
| 49 | Qwen2.5-72B (DI) | DeepInfra | 86% | $0.350 | $0.400 | $0.365 | 128,000 |
| 50 | Llama-3.3-70B (Cerebras) | Cerebras | 86% | $0.590 | $0.990 | $0.710 | 128,000 |
| 51 | Qwen2.5-72B (Hyp) | Hyperbolic | 86% | $0.400 | $0.400 | $0.400 | 128,000 |
| 52 | Qwen2.5-72B-Instruct | Qwen | 86% | $1.20 | $1.20 | $1.20 | 128,000 |
| 53 | Claude Sonnet 4.5 | Anthropic | 85% | $3.00 | $15.00 | $6.60 | 200,000 |
| 54 | o4-mini | OpenAI | 85% | $1.10 | $4.40 | $2.09 | 200,000 |
| 55 | Llama 4 Maverick | Meta | 85% | $0.200 | $0.600 | $0.320 | 128,000 |
| 56 | Command A | Cohere | 85% | $2.50 | $10.00 | $4.75 | 256,000 |
| 57 | Llama-3.1-Nemotron-70B | Nvidia | 85% | $0.350 | $0.400 | $0.365 | 128,000 |
| 58 | Claude Sonnet 3.7 | Anthropic | 84% | $3.00 | $15.00 | $6.60 | 200,000 |
| 59 | Gemini 2.5 Flash | 84% | $0.150 | $0.600 | $0.285 | 1,000,000 | |
| 60 | DeepSeek-R1-Distill-70B | DeepSeek | 84% | $0.350 | $0.880 | $0.509 | 128,000 |
About these scores
- Multitask academic knowledge across 57 subjects.
- Scores are as published by the model’s provider or a public leaderboard. Provider-reported figures are self-reported and are not independently re-run by us.
- Where a hosted variant of an open-weights model has no score of its own, it inherits the base model’s — the weights are the same.
- Prices beside each score are the provider’s published token rates. How the blended rate is calculated.
Showing the top 60 of 200 models we hold a MMLU score for.
Frequently asked questions
What is a good MMLU score?
Frontier models now sit in the high eighties and above, and the spread between them is only a few points. Because the benchmark is close to saturated at that level, MMLU is better used as a floor check (confirming broad competence) than as a way to separate the leading models from each other.
Is MMLU still a useful benchmark?
For breadth, yes; for ranking frontier models, less so. Its multiple-choice format rewards recall over reasoning, and top scores are bunched. GPQA Diamond and SWE-bench Verified separate the strongest models far more clearly.
Other leaderboards
Rankings that weigh price too
Leaderboards are recomputed on every deploy from the catalogue. Prices checked 12 August 2026.