Claude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/MClaude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/M

GPQA Leaderboard

Graduate-level science questions, "Google-proof". · Prices checked 12 August 2026

GPQA Diamond is a set of graduate-level questions in biology, physics and chemistry, written so that searching the web does not help. It isolates reasoning from recall, which makes it the cleanest signal we track for hard analytical work.

Reasoning models tend to lead this table, and they also tend to be the most expensive per token, because the internal chain of thought they generate before answering is billed as output.

Leaderboard

60 of 130 scored

Claude Fable 5 leads on GPQA Diamond at 95%. The median across the 130 models we hold a score for is 51%. The cheapest model on this board is Phi-4 at $0.133/M blended, scoring 56%.

#ModelProviderScoreInput /1MOutput /1MBlended 70/30Context
1Claude Fable 5Anthropic95%$10.00$50.00$22.001,000,000
2GPT-5.6 SolOpenAI94.6%$5.00$30.00$12.501,050,000
3Claude Opus 4.8Anthropic94%$5.00$25.00$11.001,000,000
4GPT-5.6 TerraOpenAI92.9%$2.00$12.00$5.001,050,000
5GPT-5.6 LunaOpenAI92.3%$0.200$1.20$0.5001,050,000
6o3OpenAI88%$10.00$40.00$19.00200,000
7GPT-5OpenAI84%$1.25$10.00$3.88400,000
8Gemini 2.5 ProGoogle84%$1.25$10.00$3.881,000,000
9o4-miniOpenAI81%$1.10$4.40$2.09200,000
10Claude Opus 4.6Anthropic79%$5.00$25.00$11.001,000,000
11o3-miniOpenAI79%$1.10$4.40$2.09200,000
12o1OpenAI78%$15.00$60.00$28.50200,000
13Claude Opus 4.5Anthropic76%$15.00$75.00$33.00200,000
14Claude Sonnet 4.6Anthropic75%$3.00$15.00$6.601,000,000
15Grok 3xAI75%$3.00$15.00$6.60131,000
16Grok 3 FastxAI75%$5.00$25.00$11.00131,000
17Claude Sonnet 4.5Anthropic73%$3.00$15.00$6.60200,000
18DeepSeek-R1-0528DeepSeek73%$0.550$2.19$1.0464,000
19GPT-4.1OpenAI71%$2.00$8.00$3.801,000,000
20DeepSeek-R1DeepSeek71%$0.550$2.19$1.0464,000
21DeepSeek-R1 (Groq)Groq71%$0.750$0.990$0.822128,000
22DeepSeek-R1 (Together)Together AI71%$3.00$7.00$4.2064,000
23DeepSeek-R1 (fw)Fireworks AI71%$3.00$8.00$4.5064,000
24DeepSeek-R1 (DI)DeepInfra71%$0.550$2.19$1.0464,000
25DeepSeek-R1 (Cerebras)Cerebras71%$0.550$0.990$0.68264,000
26Claude Sonnet 3.7Anthropic70%$3.00$15.00$6.60200,000
27Gemini 2.5 FlashGoogle70%$0.150$0.600$0.2851,000,000
28Sonar Reasoning ProPerplexity70%$2.00$8.00$3.80127,000
29Kimi k1.5Moonshot AI70%$2.50$10.00$4.75128,000
30Llama 4 MaverickMeta67%$0.200$0.600$0.320128,000
31Claude Sonnet 3.5Anthropic65%$3.00$15.00$6.60200,000
32DeepSeek-R1-Distill-70BDeepSeek65%$0.350$0.880$0.509128,000
33Sonar Deep ResearchPerplexity65%$2.00$8.00$3.80127,000
34QwQ-32BQwen65%$0.600$2.40$1.14131,000
35Grok 3 minixAI64%$0.300$0.500$0.360131,000
36Gemini 2.0 FlashGoogle62%$0.100$0.400$0.1901,000,000
37DeepSeek-R1-Distill-32BDeepSeek62%$0.200$0.500$0.290128,000
38GPT-4.1 miniOpenAI60%$0.400$1.60$0.7601,000,000
39DeepSeek-V3-0324DeepSeek60%$0.270$1.10$0.519128,000
40Nova PremierAmazon60%$2.50$12.50$5.50300,000
41Sonar ProPerplexity60%$3.00$15.00$6.60200,000
42Sonar ReasoningPerplexity60%$1.00$5.00$2.20127,000
43Llama-3.1-Nemotron-Ultra-253BNvidia60%$1.60$1.60$1.60128,000
44Qwen2.5-MaxQwen60%$1.60$6.40$3.0432,000
45Gemini 1.5 ProGoogle59%$1.25$5.00$2.382,000,000
46DeepSeek-V3DeepSeek59%$0.270$1.10$0.51964,000
47Llama 4 ScoutMeta57%$0.100$0.350$0.175512,000
48Claude Haiku 4.5Anthropic56%$0.800$4.00$1.76200,000
49Grok 2xAI56%$2.00$10.00$4.40131,000
50Grok 2 VisionxAI56%$2.00$10.00$4.4032,000
51Phi-4Microsoft56%$0.070$0.280$0.13316,000
52GPT-4oOpenAI53%$2.50$10.00$4.75128,000
53Command ACohere53%$2.50$10.00$4.75256,000
54Mistral Large 3Mistral52%$0.500$1.50$0.800256,000
55Mistral Large 2Mistral52%$2.00$6.00$3.20128,000
56Mistral-Large-2 (NIM)Nvidia52%$2.00$6.00$3.20128,000
57Gemini 2.0 Flash LiteGoogle51%$0.075$0.300$0.1421,000,000
58Gemini 1.5 FlashGoogle51%$0.075$0.300$0.1421,000,000
59DeepSeek-R1-Distill-14BDeepSeek51%$0.100$0.350$0.175128,000
60Llama 3.1 405BMeta51%$0.800$0.800$0.800128,000

About these scores

  • Graduate-level science questions, "Google-proof".
  • Scores are as published by the model’s provider or a public leaderboard. Provider-reported figures are self-reported and are not independently re-run by us.
  • Where a hosted variant of an open-weights model has no score of its own, it inherits the base model’s — the weights are the same.
  • Prices beside each score are the provider’s published token rates. How the blended rate is calculated.

Showing the top 60 of 130 models we hold a GPQA Diamond score for.

Frequently asked questions

What does GPQA measure?

Multi-step scientific reasoning at graduate level. The questions are deliberately "Google-proof": domain experts answer roughly two-thirds correctly even with unrestricted web access, so a high score reflects genuine reasoning rather than retrieval or memorisation.

Why do reasoning models score higher on GPQA?

They are trained to work through a problem step by step before answering, which is exactly what these questions require. That extra deliberation costs output tokens, so the models at the top of this table are usually also near the top of the price sheet.

Other leaderboards

Rankings that weigh price too

Leaderboards are recomputed on every deploy from the catalogue. Prices checked 12 August 2026.