Claude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/MClaude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/M

MATH Leaderboard

High-school competition math problems. · Prices checked 12 August 2026

The MATH benchmark uses high-school competition problems that need multi-step symbolic working rather than recall. It separates models that can hold a chain of reasoning together from those that pattern-match toward an answer.

Leaderboard

60 of 154 scored

Claude Fable 5 leads on MATH at 96%. The median across the 154 models we hold a score for is 73%. The cheapest model on this board is Phi-4 at $0.133/M blended, scoring 80%.

#ModelProviderScoreInput /1MOutput /1MBlended 70/30Context
1Claude Fable 5Anthropic96%$10.00$50.00$22.001,000,000
2o3OpenAI96%$10.00$40.00$19.00200,000
3Kimi k1.5Moonshot AI96%$2.50$10.00$4.75128,000
4GPT-5OpenAI95%$1.25$10.00$3.88400,000
5o3-miniOpenAI95%$1.10$4.40$2.09200,000
6DeepSeek-R1DeepSeek95%$0.550$2.19$1.0464,000
7DeepSeek-R1-0528DeepSeek95%$0.550$2.19$1.0464,000
8DeepSeek-R1 (Groq)Groq95%$0.750$0.990$0.822128,000
9DeepSeek-R1 (Together)Together AI95%$3.00$7.00$4.2064,000
10DeepSeek-R1 (fw)Fireworks AI95%$3.00$8.00$4.5064,000
11DeepSeek-R1 (DI)DeepInfra95%$0.550$2.19$1.0464,000
12DeepSeek-R1 (Cerebras)Cerebras95%$0.550$0.990$0.68264,000
13o4-miniOpenAI94%$1.10$4.40$2.09200,000
14o1OpenAI94%$15.00$60.00$28.50200,000
15Gemini 2.5 ProGoogle93%$1.25$10.00$3.881,000,000
16Claude Opus 4.6Anthropic92%$5.00$25.00$11.001,000,000
17DeepSeek-R1-Distill-70BDeepSeek92%$0.350$0.880$0.509128,000
18Sonar Reasoning ProPerplexity92%$2.00$8.00$3.80127,000
19DeepSeek-V3-0324DeepSeek91%$0.270$1.10$0.519128,000
20DeepSeek-V3DeepSeek90%$0.270$1.10$0.51964,000
21DeepSeek-R1-Distill-32BDeepSeek90%$0.200$0.500$0.290128,000
22QwQ-32BQwen90%$0.600$2.40$1.14131,000
23Claude Sonnet 4.6Anthropic89%$3.00$15.00$6.601,000,000
24Grok 3xAI89%$3.00$15.00$6.60131,000
25Grok 3 FastxAI89%$5.00$25.00$11.00131,000
26Claude Opus 4.5Anthropic88%$15.00$75.00$33.00200,000
27Sonar Deep ResearchPerplexity88%$2.00$8.00$3.80127,000
28GPT-4.1OpenAI87%$2.00$8.00$3.801,000,000
29Gemini 2.5 FlashGoogle87%$0.150$0.600$0.2851,000,000
30Claude Sonnet 4.5Anthropic86%$3.00$15.00$6.60200,000
31Gemini 1.5 ProGoogle86%$1.25$5.00$2.382,000,000
32Qwen2.5-MaxQwen85%$1.60$6.40$3.0432,000
33Claude Sonnet 3.7Anthropic84%$3.00$15.00$6.60200,000
34DeepSeek-R1-Distill-14BDeepSeek84%$0.100$0.350$0.175128,000
35Llama 4 MaverickMeta84%$0.200$0.600$0.320128,000
36Nova PremierAmazon84%$2.50$12.50$5.50300,000
37Sonar ReasoningPerplexity84%$1.00$5.00$2.20127,000
38Qwen2.5-72B (Groq)Groq83%$0.790$0.790$0.790128,000
39Qwen2.5-72B (Together)Together AI83%$1.20$1.20$1.2032,000
40Qwen2.5-72B (fw)Fireworks AI83%$0.900$0.900$0.90032,000
41Qwen2.5-72B (DI)DeepInfra83%$0.350$0.400$0.365128,000
42Qwen2.5-72B (Hyp)Hyperbolic83%$0.400$0.400$0.400128,000
43Qwen2.5-72B-InstructQwen83%$1.20$1.20$1.20128,000
44Qwen2.5-32B-InstructQwen83%$0.700$0.700$0.700128,000
45Grok 3 minixAI81%$0.300$0.500$0.360131,000
46Gemini 2.0 FlashGoogle80%$0.100$0.400$0.1901,000,000
47Sonar ProPerplexity80%$3.00$15.00$6.60200,000
48Llama-3.1-Nemotron-Ultra-253BNvidia80%$1.60$1.60$1.60128,000
49Phi-4Microsoft80%$0.070$0.280$0.13316,000
50Qwen2.5-14B-InstructQwen80%$0.350$0.350$0.350128,000
51Claude Sonnet 3.5Anthropic78%$3.00$15.00$6.60200,000
52GPT-4.1 miniOpenAI78%$0.400$1.60$0.7601,000,000
53Gemini 1.5 FlashGoogle78%$0.075$0.300$0.1421,000,000
54Llama 4 ScoutMeta78%$0.100$0.350$0.175512,000
55MiniMax-Text-01MiniMax78%$0.700$2.80$1.331,000,000
56Llama 3.3 70BMeta77%$0.230$0.400$0.281128,000
57Llama 3.3 70B (Groq)Groq77%$0.590$0.790$0.650128,000
58Llama 3.3 70B (Together)Together AI77%$0.880$0.880$0.880128,000
59Llama 3.3 70B (DI)DeepInfra77%$0.230$0.400$0.281128,000
60Llama-3.3-70B (Cerebras)Cerebras77%$0.590$0.990$0.710128,000

About these scores

  • High-school competition math problems.
  • Scores are as published by the model’s provider or a public leaderboard. Provider-reported figures are self-reported and are not independently re-run by us.
  • Where a hosted variant of an open-weights model has no score of its own, it inherits the base model’s — the weights are the same.
  • Prices beside each score are the provider’s published token rates. How the blended rate is calculated.

Showing the top 60 of 154 models we hold a MATH score for.

Frequently asked questions

What is the MATH benchmark?

A set of competition mathematics problems spanning algebra, geometry, number theory and precalculus. Each requires a worked solution rather than a lookup, and grading checks the final answer, so partial reasoning that lands on the wrong result scores nothing.

Other leaderboards

Rankings that weigh price too

Leaderboards are recomputed on every deploy from the catalogue. Prices checked 12 August 2026.