Claude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/MClaude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/M

Best AI Models for Math

Ranked from provider-published pricing · Prices checked 12 August 2026

The MATH benchmark uses high-school competition problems that require multi-step symbolic reasoning rather than recall. It separates models that can carry a chain of reasoning from those that pattern-match an answer.

The ranking

top 20 of 154

Claude Fable 5 from Anthropic leads this ranking at 96%. The median across the 154 qualifying models is 73%.

#ModelProviderMATH Benchmark ScoreGPQA DiamondInput /1MOutput /1MBlended 70/30
1Claude Fable 5Anthropic96%95$10.00$50.00$22.00
2o3OpenAI96%88$10.00$40.00$19.00
3Kimi k1.5Moonshot AI96%70$2.50$10.00$4.75
4GPT-5OpenAI95%84$1.25$10.00$3.88
5o3-miniOpenAI95%79$1.10$4.40$2.09
6DeepSeek-R1DeepSeek95%71$0.550$2.19$1.04
7DeepSeek-R1-0528DeepSeek95%73$0.550$2.19$1.04
8DeepSeek-R1 (Groq)Groq95%71$0.750$0.990$0.822
9DeepSeek-R1 (Together)Together AI95%71$3.00$7.00$4.20
10DeepSeek-R1 (fw)Fireworks AI95%71$3.00$8.00$4.50
11DeepSeek-R1 (DI)DeepInfra95%71$0.550$2.19$1.04
12DeepSeek-R1 (Cerebras)Cerebras95%71$0.550$0.990$0.682
13o4-miniOpenAI94%81$1.10$4.40$2.09
14o1OpenAI94%78$15.00$60.00$28.50
15Gemini 2.5 ProGoogle93%84$1.25$10.00$3.88
16Claude Opus 4.6Anthropic92%79$5.00$25.00$11.00
17DeepSeek-R1-Distill-70BDeepSeek92%65$0.350$0.880$0.509
18Sonar Reasoning ProPerplexity92%70$2.00$8.00$3.80
19DeepSeek-V3-0324DeepSeek91%60$0.270$1.10$0.519
20DeepSeek-V3DeepSeek90%59$0.270$1.10$0.519

How this list is built

  • Ranked by MATH score, highest first. Scores are as published by the model’s provider or a public leaderboard. Where a hosted variant has no score of its own, it inherits the base model’s.
  • Includes models marked available or preview; retired and deprecated SKUs are excluded.
  • Prices are the rates each provider publishes, not estimates. How the blended rate is calculated.

132 otherwise-eligible models were left out because we hold no published value for the figure this list ranks by. They are excluded rather than assumed.

Frequently asked questions

Are AI models reliable at maths?

The strongest models score highly on competition-style problems, but a benchmark score is an average, not a guarantee, a model that solves 90% of MATH problems still fails one in ten. For anything where correctness matters, have the model show its working, or verify the result with a calculator or symbolic tool rather than trusting the answer directly.

Related rankings

Price it for your workload

Rankings are recomputed on every deploy from the catalogue. Prices checked 12 August 2026.