Claude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/MClaude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/M

Best Reasoning AI Models

Ranked from provider-published pricing · Prices checked 12 August 2026

GPQA Diamond is a set of graduate-level science questions written to be "Google-proof"; retrieval does not help, so the score reflects reasoning rather than recall. It is the cleanest public signal for hard analytical work.

Reasoning models generate long internal chains of thought before answering, which you pay for as output tokens. The output rate beside each score is therefore the number to watch.

The ranking

top 20 of 130

Claude Fable 5 from Anthropic leads this ranking at 95%. The median across the 130 qualifying models is 51%.

#ModelProviderGPQA Diamond ScoreMMLUInput /1MOutput /1MBlended 70/30
1Claude Fable 5Anthropic95%90$10.00$50.00$22.00
2GPT-5.6 SolOpenAI94.6%$5.00$30.00$12.50
3Claude Opus 4.8Anthropic94%$5.00$25.00$11.00
4GPT-5.6 TerraOpenAI92.9%$2.00$12.00$5.00
5GPT-5.6 LunaOpenAI92.3%$0.200$1.20$0.500
6o3OpenAI88%90$10.00$40.00$19.00
7GPT-5OpenAI84%91$1.25$10.00$3.88
8Gemini 2.5 ProGoogle84%89$1.25$10.00$3.88
9o4-miniOpenAI81%85$1.10$4.40$2.09
10Claude Opus 4.6Anthropic79%88$5.00$25.00$11.00
11o3-miniOpenAI79%86$1.10$4.40$2.09
12o1OpenAI78%91$15.00$60.00$28.50
13Claude Opus 4.5Anthropic76%87$15.00$75.00$33.00
14Claude Sonnet 4.6Anthropic75%86$3.00$15.00$6.60
15Grok 3xAI75%86$3.00$15.00$6.60
16Grok 3 FastxAI75%86$5.00$25.00$11.00
17Claude Sonnet 4.5Anthropic73%85$3.00$15.00$6.60
18DeepSeek-R1-0528DeepSeek73%90$0.550$2.19$1.04
19GPT-4.1OpenAI71%89$2.00$8.00$3.80
20DeepSeek-R1DeepSeek71%90$0.550$2.19$1.04

How this list is built

  • Ranked by GPQA Diamond score, highest first. Scores are as published by the model’s provider or a public leaderboard. Where a hosted variant has no score of its own, it inherits the base model’s.
  • Includes models marked available or preview; retired and deprecated SKUs are excluded.
  • Prices are the rates each provider publishes, not estimates. How the blended rate is calculated.

156 otherwise-eligible models were left out because we hold no published value for the figure this list ranks by. They are excluded rather than assumed.

Frequently asked questions

What is a reasoning model?

A model trained to work through a problem step by step before committing to an answer, generating intermediate reasoning that is often hidden from the final response. The approach improves accuracy on hard multi-step problems and costs more, because all of that intermediate thinking is billed as output tokens.

Are reasoning models worth the extra cost?

For genuinely hard problems (multi-step analysis, difficult debugging, research questions) the accuracy gain usually justifies it. For classification, extraction, formatting or routine drafting they mostly burn extra output tokens without changing the answer, so routing only the hard requests to a reasoning model is the pattern that pays.

Related rankings

Price it for your workload

Rankings are recomputed on every deploy from the catalogue. Prices checked 12 August 2026.