Why AI Benchmark Scores Do not Compare
AI pricing guide · updated 2026-08-06
A benchmark score is only meaningful against the same benchmark, in the same variant, measured the same way. That sounds obvious, and it is routinely violated, by vendors, by comparison sites, and by anyone assembling a table from press releases. The result is rankings that quietly reward whoever chose the easiest test.
Variants are different benchmarks wearing the same name
SWE-bench is the clearest example. SWE-bench Verified is a human-validated subset of real GitHub issues, checked to be solvable and correctly specified. SWE-bench Pro is a deliberately harder set. They share a family name and nothing else that matters: a model scoring 64% on Pro may well be stronger than one scoring 70% on Verified.
Put those two numbers in the same column and you have inverted the ranking. This is not a hypothetical, vendors increasingly report the harder variant precisely because it differentiates, which means the more capable models are the ones most likely to look worse in a careless table.
The same applies to versioned suites. Terminal-Bench 2.0 is not Terminal-Bench 1.0. A tiered benchmark like FrontierMath reports separate figures per tier, and Tier 4 numbers are far lower than Tier 1–3 numbers for the same model.
Saturation makes old benchmarks stop discriminating
MMLU was the standard general-capability test for years. Frontier models now cluster within a few points of each other near the top, which means the gaps you can still see are mostly noise from evaluation setup rather than real capability differences.
This is why labs keep introducing new benchmarks: once everyone scores in the high eighties, the test has stopped doing its job. It also means a saturated benchmark is best read as a floor check (confirming broad competence) rather than as a way to rank the leaders.
Most published scores are self-reported
The majority of benchmark figures in circulation come from the model vendor, published in a launch post or system card, and are not independently re-run. That does not make them false. It does mean the vendor chose which benchmarks to publish, chose the prompting and scaffolding, and had every incentive to present the configuration that looked best.
Independent leaderboards help, and crowd-sourced preference ratings like LMArena Elo help more, because they are vendor-neutral and cover many models on identical terms. But even there, differences in evaluation harness, prompt formatting and retry policy move scores by several points.
The practical rule: know who reported a number and when. A score without provenance is not evidence.
Third-party aggregators contradict each other
When a model launches, comparison sites race to publish its scores, often before the vendor has released full detail. The results diverge sharply. For one recent frontier model, two aggregators reported SWE-bench Verified figures of 72.5% and 97%, a gap far too large to be a rounding difference, and a third source placed its graduate-reasoning score below the model it had just replaced.
At least one of those numbers is wrong, and from the outside there is no way to tell which. This is why our benchmark data comes only from vendor system cards and named independent leaderboards, with the source recorded, and why a model with no sourceable score shows no score at all.
How to read a benchmark table safely
Check that every model in a column was measured on the identical benchmark and variant. If a cell is empty, find out whether that means the model scored badly or simply was not tested; those are very different facts, and most tables do not distinguish them.
Treat a single score as one data point about one narrow task, not as a capability rating. Benchmarks say nothing about latency, context window, tool-calling reliability, output formatting, rate limits or price, and those frequently decide which model you should actually use.
Then weight by your own workload. A model that leads on competition mathematics may be irrelevant to a classification pipeline, and the cheapest model that clears your quality bar is usually the right answer rather than the highest scorer.
How this site handles it
We keep two tiers. A small comparable core (MMLU, GPQA Diamond, MATH, HumanEval, SWE-bench Verified and LMArena Elo) is the only data allowed to drive leaderboards, rankings and head-to-head verdicts, because those require like-for-like measurement across many models.
Everything else a vendor publishes is recorded on that model’s own page as "also reported", with the benchmark's full name including its variant, who published it, whether it was self-reported, and the date. Those figures never enter a ranking.
A benchmark is promoted into the comparable core only once enough models publish it for a genuine like-for-like table to exist. Until then, showing it as though it were comparable would mislead more than it informs.
Frequently asked questions
Can I compare SWE-bench Verified and SWE-bench Pro scores?
No. Pro is a harder set than Verified, so the two produce systematically different numbers for the same model. Comparing across them can invert the ranking, a model with a lower Pro score may be more capable than one with a higher Verified score. Only compare scores measured on the identical variant.
Are vendor-reported benchmark scores trustworthy?
Treat them as directional rather than authoritative. They are usually accurate for the configuration tested, but the vendor chose the benchmark, the prompting and the scaffolding, and did not have the result independently reproduced. Vendor-neutral leaderboards are a useful cross-check, and knowing who reported a figure matters as much as the figure.
Which AI benchmark should I actually pay attention to?
It depends what you are buying the model for: SWE-bench Verified for autonomous coding, GPQA Diamond for hard reasoning, LMArena Elo for general helpfulness as people perceive it. MMLU is close to saturated and works better as a floor check. For most decisions, price and context window matter more than any benchmark, because those are directly comparable and directly affect your bill.
Why do some models show no benchmark scores here?
Because we could not source a score from a vendor system card or a named independent leaderboard. Frontier models frequently ship months before comparable scores are published, and third-party aggregators disagree too much to be relied on. We leave the field empty rather than fill it with a number we cannot stand behind.
See it in the data
Related guides
Terms used in this guide
Published by Tokenando. Last updated 2026-08-06. Figures in this guide are computed from our own pricing index and dated where they can move; see the methodology and corrections policy.