MATH Leaderboard
High-school competition math problems. · Prices checked 12 August 2026
The MATH benchmark uses high-school competition problems that need multi-step symbolic working rather than recall. It separates models that can hold a chain of reasoning together from those that pattern-match toward an answer.
Leaderboard
60 of 154 scoredClaude Fable 5 leads on MATH at 96%. The median across the 154 models we hold a score for is 73%. The cheapest model on this board is Phi-4 at $0.133/M blended, scoring 80%.
| # | Model | Provider | Score | Input /1M | Output /1M | Blended 70/30 | Context |
|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | Anthropic | 96% | $10.00 | $50.00 | $22.00 | 1,000,000 |
| 2 | o3 | OpenAI | 96% | $10.00 | $40.00 | $19.00 | 200,000 |
| 3 | Kimi k1.5 | Moonshot AI | 96% | $2.50 | $10.00 | $4.75 | 128,000 |
| 4 | GPT-5 | OpenAI | 95% | $1.25 | $10.00 | $3.88 | 400,000 |
| 5 | o3-mini | OpenAI | 95% | $1.10 | $4.40 | $2.09 | 200,000 |
| 6 | DeepSeek-R1 | DeepSeek | 95% | $0.550 | $2.19 | $1.04 | 64,000 |
| 7 | DeepSeek-R1-0528 | DeepSeek | 95% | $0.550 | $2.19 | $1.04 | 64,000 |
| 8 | DeepSeek-R1 (Groq) | Groq | 95% | $0.750 | $0.990 | $0.822 | 128,000 |
| 9 | DeepSeek-R1 (Together) | Together AI | 95% | $3.00 | $7.00 | $4.20 | 64,000 |
| 10 | DeepSeek-R1 (fw) | Fireworks AI | 95% | $3.00 | $8.00 | $4.50 | 64,000 |
| 11 | DeepSeek-R1 (DI) | DeepInfra | 95% | $0.550 | $2.19 | $1.04 | 64,000 |
| 12 | DeepSeek-R1 (Cerebras) | Cerebras | 95% | $0.550 | $0.990 | $0.682 | 64,000 |
| 13 | o4-mini | OpenAI | 94% | $1.10 | $4.40 | $2.09 | 200,000 |
| 14 | o1 | OpenAI | 94% | $15.00 | $60.00 | $28.50 | 200,000 |
| 15 | Gemini 2.5 Pro | 93% | $1.25 | $10.00 | $3.88 | 1,000,000 | |
| 16 | Claude Opus 4.6 | Anthropic | 92% | $5.00 | $25.00 | $11.00 | 1,000,000 |
| 17 | DeepSeek-R1-Distill-70B | DeepSeek | 92% | $0.350 | $0.880 | $0.509 | 128,000 |
| 18 | Sonar Reasoning Pro | Perplexity | 92% | $2.00 | $8.00 | $3.80 | 127,000 |
| 19 | DeepSeek-V3-0324 | DeepSeek | 91% | $0.270 | $1.10 | $0.519 | 128,000 |
| 20 | DeepSeek-V3 | DeepSeek | 90% | $0.270 | $1.10 | $0.519 | 64,000 |
| 21 | DeepSeek-R1-Distill-32B | DeepSeek | 90% | $0.200 | $0.500 | $0.290 | 128,000 |
| 22 | QwQ-32B | Qwen | 90% | $0.600 | $2.40 | $1.14 | 131,000 |
| 23 | Claude Sonnet 4.6 | Anthropic | 89% | $3.00 | $15.00 | $6.60 | 1,000,000 |
| 24 | Grok 3 | xAI | 89% | $3.00 | $15.00 | $6.60 | 131,000 |
| 25 | Grok 3 Fast | xAI | 89% | $5.00 | $25.00 | $11.00 | 131,000 |
| 26 | Claude Opus 4.5 | Anthropic | 88% | $15.00 | $75.00 | $33.00 | 200,000 |
| 27 | Sonar Deep Research | Perplexity | 88% | $2.00 | $8.00 | $3.80 | 127,000 |
| 28 | GPT-4.1 | OpenAI | 87% | $2.00 | $8.00 | $3.80 | 1,000,000 |
| 29 | Gemini 2.5 Flash | 87% | $0.150 | $0.600 | $0.285 | 1,000,000 | |
| 30 | Claude Sonnet 4.5 | Anthropic | 86% | $3.00 | $15.00 | $6.60 | 200,000 |
| 31 | Gemini 1.5 Pro | 86% | $1.25 | $5.00 | $2.38 | 2,000,000 | |
| 32 | Qwen2.5-Max | Qwen | 85% | $1.60 | $6.40 | $3.04 | 32,000 |
| 33 | Claude Sonnet 3.7 | Anthropic | 84% | $3.00 | $15.00 | $6.60 | 200,000 |
| 34 | DeepSeek-R1-Distill-14B | DeepSeek | 84% | $0.100 | $0.350 | $0.175 | 128,000 |
| 35 | Llama 4 Maverick | Meta | 84% | $0.200 | $0.600 | $0.320 | 128,000 |
| 36 | Nova Premier | Amazon | 84% | $2.50 | $12.50 | $5.50 | 300,000 |
| 37 | Sonar Reasoning | Perplexity | 84% | $1.00 | $5.00 | $2.20 | 127,000 |
| 38 | Qwen2.5-72B (Groq) | Groq | 83% | $0.790 | $0.790 | $0.790 | 128,000 |
| 39 | Qwen2.5-72B (Together) | Together AI | 83% | $1.20 | $1.20 | $1.20 | 32,000 |
| 40 | Qwen2.5-72B (fw) | Fireworks AI | 83% | $0.900 | $0.900 | $0.900 | 32,000 |
| 41 | Qwen2.5-72B (DI) | DeepInfra | 83% | $0.350 | $0.400 | $0.365 | 128,000 |
| 42 | Qwen2.5-72B (Hyp) | Hyperbolic | 83% | $0.400 | $0.400 | $0.400 | 128,000 |
| 43 | Qwen2.5-72B-Instruct | Qwen | 83% | $1.20 | $1.20 | $1.20 | 128,000 |
| 44 | Qwen2.5-32B-Instruct | Qwen | 83% | $0.700 | $0.700 | $0.700 | 128,000 |
| 45 | Grok 3 mini | xAI | 81% | $0.300 | $0.500 | $0.360 | 131,000 |
| 46 | Gemini 2.0 Flash | 80% | $0.100 | $0.400 | $0.190 | 1,000,000 | |
| 47 | Sonar Pro | Perplexity | 80% | $3.00 | $15.00 | $6.60 | 200,000 |
| 48 | Llama-3.1-Nemotron-Ultra-253B | Nvidia | 80% | $1.60 | $1.60 | $1.60 | 128,000 |
| 49 | Phi-4 | Microsoft | 80% | $0.070 | $0.280 | $0.133 | 16,000 |
| 50 | Qwen2.5-14B-Instruct | Qwen | 80% | $0.350 | $0.350 | $0.350 | 128,000 |
| 51 | Claude Sonnet 3.5 | Anthropic | 78% | $3.00 | $15.00 | $6.60 | 200,000 |
| 52 | GPT-4.1 mini | OpenAI | 78% | $0.400 | $1.60 | $0.760 | 1,000,000 |
| 53 | Gemini 1.5 Flash | 78% | $0.075 | $0.300 | $0.142 | 1,000,000 | |
| 54 | Llama 4 Scout | Meta | 78% | $0.100 | $0.350 | $0.175 | 512,000 |
| 55 | MiniMax-Text-01 | MiniMax | 78% | $0.700 | $2.80 | $1.33 | 1,000,000 |
| 56 | Llama 3.3 70B | Meta | 77% | $0.230 | $0.400 | $0.281 | 128,000 |
| 57 | Llama 3.3 70B (Groq) | Groq | 77% | $0.590 | $0.790 | $0.650 | 128,000 |
| 58 | Llama 3.3 70B (Together) | Together AI | 77% | $0.880 | $0.880 | $0.880 | 128,000 |
| 59 | Llama 3.3 70B (DI) | DeepInfra | 77% | $0.230 | $0.400 | $0.281 | 128,000 |
| 60 | Llama-3.3-70B (Cerebras) | Cerebras | 77% | $0.590 | $0.990 | $0.710 | 128,000 |
About these scores
- High-school competition math problems.
- Scores are as published by the model’s provider or a public leaderboard. Provider-reported figures are self-reported and are not independently re-run by us.
- Where a hosted variant of an open-weights model has no score of its own, it inherits the base model’s — the weights are the same.
- Prices beside each score are the provider’s published token rates. How the blended rate is calculated.
Showing the top 60 of 154 models we hold a MATH score for.
Frequently asked questions
What is the MATH benchmark?
A set of competition mathematics problems spanning algebra, geometry, number theory and precalculus. Each requires a worked solution rather than a lookup, and grading checks the final answer, so partial reasoning that lands on the wrong result scores nothing.
Other leaderboards
Rankings that weigh price too
Leaderboards are recomputed on every deploy from the catalogue. Prices checked 12 August 2026.