Benchmarks · Capability vs cost
AI Benchmarks
6 leaderboardsSix benchmark leaderboards, each showing every model’s score beside the price you pay to run it — the trade-off most leaderboards leave out. For the raw rates see the AI model pricing index.
Leaderboards
6Frequently asked questions
Which AI benchmark matters most?
It depends what you are buying the model for. SWE-bench Verified is the best signal for autonomous coding, GPQA Diamond for hard reasoning, and LMArena Elo for general helpfulness as people actually perceive it. MMLU is now close to saturated at the top and works better as a breadth check than a way to rank leading models.
Are AI benchmark scores reliable?
Treat them as directional. Most published figures are self-reported by the provider and not independently re-run, evaluation harnesses differ between labs, and a single number averages over a fixed set of problems that may not resemble your workload. They are useful for shortlisting, not for final selection.
Does a higher benchmark score justify a higher price?
Sometimes, and these tables are built so you can judge that directly: every leaderboard shows each model’s token price beside its score. The strongest models often cost many times more per token than models a few points behind them, which is worth it for hard problems and wasteful for routine ones.
Rankings that weigh price too
Scores are provider-published or from public leaderboards, not independently re-run. Prices checked 12 August 2026.