Claude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/MClaude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/M

Benchmarks · Capability vs cost

AI Benchmarks

6 leaderboards

Six benchmark leaderboards, each showing every model’s score beside the price you pay to run it — the trade-off most leaderboards leave out. For the raw rates see the AI model pricing index.

Leaderboards

6

Frequently asked questions

Which AI benchmark matters most?

It depends what you are buying the model for. SWE-bench Verified is the best signal for autonomous coding, GPQA Diamond for hard reasoning, and LMArena Elo for general helpfulness as people actually perceive it. MMLU is now close to saturated at the top and works better as a breadth check than a way to rank leading models.

Are AI benchmark scores reliable?

Treat them as directional. Most published figures are self-reported by the provider and not independently re-run, evaluation harnesses differ between labs, and a single number averages over a fixed set of problems that may not resemble your workload. They are useful for shortlisting, not for final selection.

Does a higher benchmark score justify a higher price?

Sometimes, and these tables are built so you can judge that directly: every leaderboard shows each model’s token price beside its score. The strongest models often cost many times more per token than models a few points behind them, which is worth it for hard problems and wasteful for routine ones.

Rankings that weigh price too

Scores are provider-published or from public leaderboards, not independently re-run. Prices checked 12 August 2026.