GPQA Leaderboard
Graduate-level science questions, "Google-proof". · Prices checked 12 August 2026
GPQA Diamond is a set of graduate-level questions in biology, physics and chemistry, written so that searching the web does not help. It isolates reasoning from recall, which makes it the cleanest signal we track for hard analytical work.
Reasoning models tend to lead this table, and they also tend to be the most expensive per token, because the internal chain of thought they generate before answering is billed as output.
Leaderboard
60 of 130 scoredClaude Fable 5 leads on GPQA Diamond at 95%. The median across the 130 models we hold a score for is 51%. The cheapest model on this board is Phi-4 at $0.133/M blended, scoring 56%.
| # | Model | Provider | Score | Input /1M | Output /1M | Blended 70/30 | Context |
|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | Anthropic | 95% | $10.00 | $50.00 | $22.00 | 1,000,000 |
| 2 | GPT-5.6 Sol | OpenAI | 94.6% | $5.00 | $30.00 | $12.50 | 1,050,000 |
| 3 | Claude Opus 4.8 | Anthropic | 94% | $5.00 | $25.00 | $11.00 | 1,000,000 |
| 4 | GPT-5.6 Terra | OpenAI | 92.9% | $2.00 | $12.00 | $5.00 | 1,050,000 |
| 5 | GPT-5.6 Luna | OpenAI | 92.3% | $0.200 | $1.20 | $0.500 | 1,050,000 |
| 6 | o3 | OpenAI | 88% | $10.00 | $40.00 | $19.00 | 200,000 |
| 7 | GPT-5 | OpenAI | 84% | $1.25 | $10.00 | $3.88 | 400,000 |
| 8 | Gemini 2.5 Pro | 84% | $1.25 | $10.00 | $3.88 | 1,000,000 | |
| 9 | o4-mini | OpenAI | 81% | $1.10 | $4.40 | $2.09 | 200,000 |
| 10 | Claude Opus 4.6 | Anthropic | 79% | $5.00 | $25.00 | $11.00 | 1,000,000 |
| 11 | o3-mini | OpenAI | 79% | $1.10 | $4.40 | $2.09 | 200,000 |
| 12 | o1 | OpenAI | 78% | $15.00 | $60.00 | $28.50 | 200,000 |
| 13 | Claude Opus 4.5 | Anthropic | 76% | $15.00 | $75.00 | $33.00 | 200,000 |
| 14 | Claude Sonnet 4.6 | Anthropic | 75% | $3.00 | $15.00 | $6.60 | 1,000,000 |
| 15 | Grok 3 | xAI | 75% | $3.00 | $15.00 | $6.60 | 131,000 |
| 16 | Grok 3 Fast | xAI | 75% | $5.00 | $25.00 | $11.00 | 131,000 |
| 17 | Claude Sonnet 4.5 | Anthropic | 73% | $3.00 | $15.00 | $6.60 | 200,000 |
| 18 | DeepSeek-R1-0528 | DeepSeek | 73% | $0.550 | $2.19 | $1.04 | 64,000 |
| 19 | GPT-4.1 | OpenAI | 71% | $2.00 | $8.00 | $3.80 | 1,000,000 |
| 20 | DeepSeek-R1 | DeepSeek | 71% | $0.550 | $2.19 | $1.04 | 64,000 |
| 21 | DeepSeek-R1 (Groq) | Groq | 71% | $0.750 | $0.990 | $0.822 | 128,000 |
| 22 | DeepSeek-R1 (Together) | Together AI | 71% | $3.00 | $7.00 | $4.20 | 64,000 |
| 23 | DeepSeek-R1 (fw) | Fireworks AI | 71% | $3.00 | $8.00 | $4.50 | 64,000 |
| 24 | DeepSeek-R1 (DI) | DeepInfra | 71% | $0.550 | $2.19 | $1.04 | 64,000 |
| 25 | DeepSeek-R1 (Cerebras) | Cerebras | 71% | $0.550 | $0.990 | $0.682 | 64,000 |
| 26 | Claude Sonnet 3.7 | Anthropic | 70% | $3.00 | $15.00 | $6.60 | 200,000 |
| 27 | Gemini 2.5 Flash | 70% | $0.150 | $0.600 | $0.285 | 1,000,000 | |
| 28 | Sonar Reasoning Pro | Perplexity | 70% | $2.00 | $8.00 | $3.80 | 127,000 |
| 29 | Kimi k1.5 | Moonshot AI | 70% | $2.50 | $10.00 | $4.75 | 128,000 |
| 30 | Llama 4 Maverick | Meta | 67% | $0.200 | $0.600 | $0.320 | 128,000 |
| 31 | Claude Sonnet 3.5 | Anthropic | 65% | $3.00 | $15.00 | $6.60 | 200,000 |
| 32 | DeepSeek-R1-Distill-70B | DeepSeek | 65% | $0.350 | $0.880 | $0.509 | 128,000 |
| 33 | Sonar Deep Research | Perplexity | 65% | $2.00 | $8.00 | $3.80 | 127,000 |
| 34 | QwQ-32B | Qwen | 65% | $0.600 | $2.40 | $1.14 | 131,000 |
| 35 | Grok 3 mini | xAI | 64% | $0.300 | $0.500 | $0.360 | 131,000 |
| 36 | Gemini 2.0 Flash | 62% | $0.100 | $0.400 | $0.190 | 1,000,000 | |
| 37 | DeepSeek-R1-Distill-32B | DeepSeek | 62% | $0.200 | $0.500 | $0.290 | 128,000 |
| 38 | GPT-4.1 mini | OpenAI | 60% | $0.400 | $1.60 | $0.760 | 1,000,000 |
| 39 | DeepSeek-V3-0324 | DeepSeek | 60% | $0.270 | $1.10 | $0.519 | 128,000 |
| 40 | Nova Premier | Amazon | 60% | $2.50 | $12.50 | $5.50 | 300,000 |
| 41 | Sonar Pro | Perplexity | 60% | $3.00 | $15.00 | $6.60 | 200,000 |
| 42 | Sonar Reasoning | Perplexity | 60% | $1.00 | $5.00 | $2.20 | 127,000 |
| 43 | Llama-3.1-Nemotron-Ultra-253B | Nvidia | 60% | $1.60 | $1.60 | $1.60 | 128,000 |
| 44 | Qwen2.5-Max | Qwen | 60% | $1.60 | $6.40 | $3.04 | 32,000 |
| 45 | Gemini 1.5 Pro | 59% | $1.25 | $5.00 | $2.38 | 2,000,000 | |
| 46 | DeepSeek-V3 | DeepSeek | 59% | $0.270 | $1.10 | $0.519 | 64,000 |
| 47 | Llama 4 Scout | Meta | 57% | $0.100 | $0.350 | $0.175 | 512,000 |
| 48 | Claude Haiku 4.5 | Anthropic | 56% | $0.800 | $4.00 | $1.76 | 200,000 |
| 49 | Grok 2 | xAI | 56% | $2.00 | $10.00 | $4.40 | 131,000 |
| 50 | Grok 2 Vision | xAI | 56% | $2.00 | $10.00 | $4.40 | 32,000 |
| 51 | Phi-4 | Microsoft | 56% | $0.070 | $0.280 | $0.133 | 16,000 |
| 52 | GPT-4o | OpenAI | 53% | $2.50 | $10.00 | $4.75 | 128,000 |
| 53 | Command A | Cohere | 53% | $2.50 | $10.00 | $4.75 | 256,000 |
| 54 | Mistral Large 3 | Mistral | 52% | $0.500 | $1.50 | $0.800 | 256,000 |
| 55 | Mistral Large 2 | Mistral | 52% | $2.00 | $6.00 | $3.20 | 128,000 |
| 56 | Mistral-Large-2 (NIM) | Nvidia | 52% | $2.00 | $6.00 | $3.20 | 128,000 |
| 57 | Gemini 2.0 Flash Lite | 51% | $0.075 | $0.300 | $0.142 | 1,000,000 | |
| 58 | Gemini 1.5 Flash | 51% | $0.075 | $0.300 | $0.142 | 1,000,000 | |
| 59 | DeepSeek-R1-Distill-14B | DeepSeek | 51% | $0.100 | $0.350 | $0.175 | 128,000 |
| 60 | Llama 3.1 405B | Meta | 51% | $0.800 | $0.800 | $0.800 | 128,000 |
About these scores
- Graduate-level science questions, "Google-proof".
- Scores are as published by the model’s provider or a public leaderboard. Provider-reported figures are self-reported and are not independently re-run by us.
- Where a hosted variant of an open-weights model has no score of its own, it inherits the base model’s — the weights are the same.
- Prices beside each score are the provider’s published token rates. How the blended rate is calculated.
Showing the top 60 of 130 models we hold a GPQA Diamond score for.
Frequently asked questions
What does GPQA measure?
Multi-step scientific reasoning at graduate level. The questions are deliberately "Google-proof": domain experts answer roughly two-thirds correctly even with unrestricted web access, so a high score reflects genuine reasoning rather than retrieval or memorisation.
Why do reasoning models score higher on GPQA?
They are trained to work through a problem step by step before answering, which is exactly what these questions require. That extra deliberation costs output tokens, so the models at the top of this table are usually also near the top of the price sheet.
Other leaderboards
Rankings that weigh price too
Leaderboards are recomputed on every deploy from the catalogue. Prices checked 12 August 2026.