Best Reasoning AI Models
Ranked from provider-published pricing · Prices checked 12 August 2026
GPQA Diamond is a set of graduate-level science questions written to be "Google-proof"; retrieval does not help, so the score reflects reasoning rather than recall. It is the cleanest public signal for hard analytical work.
Reasoning models generate long internal chains of thought before answering, which you pay for as output tokens. The output rate beside each score is therefore the number to watch.
The ranking
top 20 of 130Claude Fable 5 from Anthropic leads this ranking at 95%. The median across the 130 qualifying models is 51%.
| # | Model | Provider | GPQA Diamond Score | MMLU | Input /1M | Output /1M | Blended 70/30 |
|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | Anthropic | 95% | 90 | $10.00 | $50.00 | $22.00 |
| 2 | GPT-5.6 Sol | OpenAI | 94.6% | — | $5.00 | $30.00 | $12.50 |
| 3 | Claude Opus 4.8 | Anthropic | 94% | — | $5.00 | $25.00 | $11.00 |
| 4 | GPT-5.6 Terra | OpenAI | 92.9% | — | $2.00 | $12.00 | $5.00 |
| 5 | GPT-5.6 Luna | OpenAI | 92.3% | — | $0.200 | $1.20 | $0.500 |
| 6 | o3 | OpenAI | 88% | 90 | $10.00 | $40.00 | $19.00 |
| 7 | GPT-5 | OpenAI | 84% | 91 | $1.25 | $10.00 | $3.88 |
| 8 | Gemini 2.5 Pro | 84% | 89 | $1.25 | $10.00 | $3.88 | |
| 9 | o4-mini | OpenAI | 81% | 85 | $1.10 | $4.40 | $2.09 |
| 10 | Claude Opus 4.6 | Anthropic | 79% | 88 | $5.00 | $25.00 | $11.00 |
| 11 | o3-mini | OpenAI | 79% | 86 | $1.10 | $4.40 | $2.09 |
| 12 | o1 | OpenAI | 78% | 91 | $15.00 | $60.00 | $28.50 |
| 13 | Claude Opus 4.5 | Anthropic | 76% | 87 | $15.00 | $75.00 | $33.00 |
| 14 | Claude Sonnet 4.6 | Anthropic | 75% | 86 | $3.00 | $15.00 | $6.60 |
| 15 | Grok 3 | xAI | 75% | 86 | $3.00 | $15.00 | $6.60 |
| 16 | Grok 3 Fast | xAI | 75% | 86 | $5.00 | $25.00 | $11.00 |
| 17 | Claude Sonnet 4.5 | Anthropic | 73% | 85 | $3.00 | $15.00 | $6.60 |
| 18 | DeepSeek-R1-0528 | DeepSeek | 73% | 90 | $0.550 | $2.19 | $1.04 |
| 19 | GPT-4.1 | OpenAI | 71% | 89 | $2.00 | $8.00 | $3.80 |
| 20 | DeepSeek-R1 | DeepSeek | 71% | 90 | $0.550 | $2.19 | $1.04 |
How this list is built
- Ranked by GPQA Diamond score, highest first. Scores are as published by the model’s provider or a public leaderboard. Where a hosted variant has no score of its own, it inherits the base model’s.
- Includes models marked available or preview; retired and deprecated SKUs are excluded.
- Prices are the rates each provider publishes, not estimates. How the blended rate is calculated.
156 otherwise-eligible models were left out because we hold no published value for the figure this list ranks by. They are excluded rather than assumed.
Frequently asked questions
What is a reasoning model?
A model trained to work through a problem step by step before committing to an answer, generating intermediate reasoning that is often hidden from the final response. The approach improves accuracy on hard multi-step problems and costs more, because all of that intermediate thinking is billed as output tokens.
Are reasoning models worth the extra cost?
For genuinely hard problems (multi-step analysis, difficult debugging, research questions) the accuracy gain usually justifies it. For classification, extraction, formatting or routine drafting they mostly burn extra output tokens without changing the answer, so routing only the hard requests to a reasoning model is the pattern that pays.
Related rankings
Price it for your workload
Rankings are recomputed on every deploy from the catalogue. Prices checked 12 August 2026.