Claude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/MClaude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/M

MMLU Leaderboard

Multitask academic knowledge across 57 subjects. · Prices checked 12 August 2026

MMLU tests multiple-choice knowledge across 57 subjects, from elementary mathematics to professional law. It was the standard general-capability benchmark for years and remains a useful breadth check.

It is now close to saturated at the top: the strongest models cluster within a few points of each other, so small differences here mean much less than they did when the benchmark was new.

Leaderboard

60 of 200 scored

GPT-5 leads on MMLU at 91%. The median across the 200 models we hold a score for is 79%. The cheapest model on this board is Llama 3.3 70B at $0.281/M blended, scoring 86%.

#ModelProviderScoreInput /1MOutput /1MBlended 70/30Context
1GPT-5OpenAI91%$1.25$10.00$3.88400,000
2o1OpenAI91%$15.00$60.00$28.50200,000
3Claude Fable 5Anthropic90%$10.00$50.00$22.001,000,000
4o3OpenAI90%$10.00$40.00$19.00200,000
5DeepSeek-R1DeepSeek90%$0.550$2.19$1.0464,000
6DeepSeek-R1-0528DeepSeek90%$0.550$2.19$1.0464,000
7DeepSeek-R1 (Groq)Groq90%$0.750$0.990$0.822128,000
8DeepSeek-R1 (Together)Together AI90%$3.00$7.00$4.2064,000
9DeepSeek-R1 (fw)Fireworks AI90%$3.00$8.00$4.5064,000
10DeepSeek-R1 (DI)DeepInfra90%$0.550$2.19$1.0464,000
11DeepSeek-R1 (Cerebras)Cerebras90%$0.550$0.990$0.68264,000
12GPT-4.1OpenAI89%$2.00$8.00$3.801,000,000
13Gemini 2.5 ProGoogle89%$1.25$10.00$3.881,000,000
14Claude Opus 4.6Anthropic88%$5.00$25.00$11.001,000,000
15GPT-4oOpenAI88%$2.50$10.00$4.75128,000
16DeepSeek-V3DeepSeek88%$0.270$1.10$0.51964,000
17DeepSeek-V3-0324DeepSeek88%$0.270$1.10$0.519128,000
18Llama 3.1 405BMeta88%$0.800$0.800$0.800128,000
19Nova PremierAmazon88%$2.50$12.50$5.50300,000
20Sonar Reasoning ProPerplexity88%$2.00$8.00$3.80127,000
21Llama 3.1 405B (Groq)Groq88%$2.99$2.99$2.99128,000
22Llama 3.1 405B (Together)Together AI88%$3.50$3.50$3.50128,000
23Llama 3.1 405B (fw)Fireworks AI88%$3.00$3.00$3.00131,000
24Llama 3.1 405B (Rep)Replicate88%$9.50$9.50$9.50128,000
25Llama 3.1 405B (DI)DeepInfra88%$0.800$0.800$0.800128,000
26Llama-3.1-405B (Cerebras)Cerebras88%$0.990$0.990$0.990128,000
27Llama 3.1 405B (Hyp)Hyperbolic88%$4.00$4.00$4.00128,000
28Llama-3.1-Nemotron-Ultra-253BNvidia88%$1.60$1.60$1.60128,000
29Kimi k1.5Moonshot AI88%$2.50$10.00$4.75128,000
30Claude Opus 4.5Anthropic87%$15.00$75.00$33.00200,000
31Sonar Deep ResearchPerplexity87%$2.00$8.00$3.80127,000
32Qwen2.5-MaxQwen87%$1.60$6.40$3.0432,000
33Claude Sonnet 4.6Anthropic86%$3.00$15.00$6.601,000,000
34GPT-4 TurboOpenAI86%$10.00$30.00$16.00128,000
35o3-miniOpenAI86%$1.10$4.40$2.09200,000
36Gemini 1.5 ProGoogle86%$1.25$5.00$2.382,000,000
37Grok 3xAI86%$3.00$15.00$6.60131,000
38Grok 3 FastxAI86%$5.00$25.00$11.00131,000
39Llama 3.3 70BMeta86%$0.230$0.400$0.281128,000
40Llama 3.2 90B VisionMeta86%$0.880$0.880$0.880128,000
41Nova ProAmazon86%$0.800$3.20$1.52300,000
42Sonar ProPerplexity86%$3.00$15.00$6.60200,000
43Llama 3.3 70B (Groq)Groq86%$0.590$0.790$0.650128,000
44Qwen2.5-72B (Groq)Groq86%$0.790$0.790$0.790128,000
45Llama 3.3 70B (Together)Together AI86%$0.880$0.880$0.880128,000
46Qwen2.5-72B (Together)Together AI86%$1.20$1.20$1.2032,000
47Qwen2.5-72B (fw)Fireworks AI86%$0.900$0.900$0.90032,000
48Llama 3.3 70B (DI)DeepInfra86%$0.230$0.400$0.281128,000
49Qwen2.5-72B (DI)DeepInfra86%$0.350$0.400$0.365128,000
50Llama-3.3-70B (Cerebras)Cerebras86%$0.590$0.990$0.710128,000
51Qwen2.5-72B (Hyp)Hyperbolic86%$0.400$0.400$0.400128,000
52Qwen2.5-72B-InstructQwen86%$1.20$1.20$1.20128,000
53Claude Sonnet 4.5Anthropic85%$3.00$15.00$6.60200,000
54o4-miniOpenAI85%$1.10$4.40$2.09200,000
55Llama 4 MaverickMeta85%$0.200$0.600$0.320128,000
56Command ACohere85%$2.50$10.00$4.75256,000
57Llama-3.1-Nemotron-70BNvidia85%$0.350$0.400$0.365128,000
58Claude Sonnet 3.7Anthropic84%$3.00$15.00$6.60200,000
59Gemini 2.5 FlashGoogle84%$0.150$0.600$0.2851,000,000
60DeepSeek-R1-Distill-70BDeepSeek84%$0.350$0.880$0.509128,000

About these scores

  • Multitask academic knowledge across 57 subjects.
  • Scores are as published by the model’s provider or a public leaderboard. Provider-reported figures are self-reported and are not independently re-run by us.
  • Where a hosted variant of an open-weights model has no score of its own, it inherits the base model’s — the weights are the same.
  • Prices beside each score are the provider’s published token rates. How the blended rate is calculated.

Showing the top 60 of 200 models we hold a MMLU score for.

Frequently asked questions

What is a good MMLU score?

Frontier models now sit in the high eighties and above, and the spread between them is only a few points. Because the benchmark is close to saturated at that level, MMLU is better used as a floor check (confirming broad competence) than as a way to separate the leading models from each other.

Is MMLU still a useful benchmark?

For breadth, yes; for ranking frontier models, less so. Its multiple-choice format rewards recall over reasoning, and top scores are bunched. GPQA Diamond and SWE-bench Verified separate the strongest models far more clearly.

Other leaderboards

Rankings that weigh price too

Leaderboards are recomputed on every deploy from the catalogue. Prices checked 12 August 2026.