Best AI Models for Math
Ranked from provider-published pricing · Prices checked 12 August 2026
The MATH benchmark uses high-school competition problems that require multi-step symbolic reasoning rather than recall. It separates models that can carry a chain of reasoning from those that pattern-match an answer.
The ranking
top 20 of 154Claude Fable 5 from Anthropic leads this ranking at 96%. The median across the 154 qualifying models is 73%.
| # | Model | Provider | MATH Benchmark Score | GPQA Diamond | Input /1M | Output /1M | Blended 70/30 |
|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | Anthropic | 96% | 95 | $10.00 | $50.00 | $22.00 |
| 2 | o3 | OpenAI | 96% | 88 | $10.00 | $40.00 | $19.00 |
| 3 | Kimi k1.5 | Moonshot AI | 96% | 70 | $2.50 | $10.00 | $4.75 |
| 4 | GPT-5 | OpenAI | 95% | 84 | $1.25 | $10.00 | $3.88 |
| 5 | o3-mini | OpenAI | 95% | 79 | $1.10 | $4.40 | $2.09 |
| 6 | DeepSeek-R1 | DeepSeek | 95% | 71 | $0.550 | $2.19 | $1.04 |
| 7 | DeepSeek-R1-0528 | DeepSeek | 95% | 73 | $0.550 | $2.19 | $1.04 |
| 8 | DeepSeek-R1 (Groq) | Groq | 95% | 71 | $0.750 | $0.990 | $0.822 |
| 9 | DeepSeek-R1 (Together) | Together AI | 95% | 71 | $3.00 | $7.00 | $4.20 |
| 10 | DeepSeek-R1 (fw) | Fireworks AI | 95% | 71 | $3.00 | $8.00 | $4.50 |
| 11 | DeepSeek-R1 (DI) | DeepInfra | 95% | 71 | $0.550 | $2.19 | $1.04 |
| 12 | DeepSeek-R1 (Cerebras) | Cerebras | 95% | 71 | $0.550 | $0.990 | $0.682 |
| 13 | o4-mini | OpenAI | 94% | 81 | $1.10 | $4.40 | $2.09 |
| 14 | o1 | OpenAI | 94% | 78 | $15.00 | $60.00 | $28.50 |
| 15 | Gemini 2.5 Pro | 93% | 84 | $1.25 | $10.00 | $3.88 | |
| 16 | Claude Opus 4.6 | Anthropic | 92% | 79 | $5.00 | $25.00 | $11.00 |
| 17 | DeepSeek-R1-Distill-70B | DeepSeek | 92% | 65 | $0.350 | $0.880 | $0.509 |
| 18 | Sonar Reasoning Pro | Perplexity | 92% | 70 | $2.00 | $8.00 | $3.80 |
| 19 | DeepSeek-V3-0324 | DeepSeek | 91% | 60 | $0.270 | $1.10 | $0.519 |
| 20 | DeepSeek-V3 | DeepSeek | 90% | 59 | $0.270 | $1.10 | $0.519 |
How this list is built
- Ranked by MATH score, highest first. Scores are as published by the model’s provider or a public leaderboard. Where a hosted variant has no score of its own, it inherits the base model’s.
- Includes models marked available or preview; retired and deprecated SKUs are excluded.
- Prices are the rates each provider publishes, not estimates. How the blended rate is calculated.
132 otherwise-eligible models were left out because we hold no published value for the figure this list ranks by. They are excluded rather than assumed.
Frequently asked questions
Are AI models reliable at maths?
The strongest models score highly on competition-style problems, but a benchmark score is an average, not a guarantee, a model that solves 90% of MATH problems still fails one in ten. For anything where correctness matters, have the model show its working, or verify the result with a calculator or symbolic tool rather than trusting the answer directly.
Related rankings
Price it for your workload
Rankings are recomputed on every deploy from the catalogue. Prices checked 12 August 2026.