Blended Cost Explained
AI pricing guide · updated 2026-08-07
Every model has two prices, which makes a straight league table impossible. A blended rate solves that by weighting the two into a single figure. It is a genuinely useful shortlisting device and a poor budgeting tool, and the difference between those two uses is worth understanding before you rank anything on it.
What the number actually is
A blended rate is a weighted average of a model's input and output prices. The default on this site weights input at 70 per cent and output at 30 per cent, so a model charging $2 input and $12 output blends to (2 × 0.7) + (12 × 0.3), which is $5.00 per million tokens.
Nothing about that figure is published by the provider. It is a calculated comparison metric, which is why every table here labels it as one. The provider publishes two numbers; the third is ours.
Why a single number is needed at all
Without one, "which model is cheapest" has no answer. Across the 271 text models we track, a model can be cheapest on input and among the dearest on output. Sorting by either column alone produces a ranking that flatters one workload shape and penalises another.
A blend picks an assumption and applies it consistently. That makes the ordering reproducible and comparable, which is exactly what a shortlist needs.
The assumption is the whole catch
70/30 describes a short prompt and a longer completion. Plenty of real workloads look nothing like that. Retrieval sends far more than it receives, often 80/20 or steeper. Agent loops and long-form generation invert it, sometimes to 20/80.
Because the median model in our catalogue charges twice as much for output as for input, and 117 of 271 charge at least three times as much, changing the weighting genuinely reorders the list. This is why our rankings publish several weightings rather than one, and why the same catalogue produces a different "cheapest" model for retrieval than for generation.
When to use it and when to stop
Use a blended rate to narrow a field of hundreds down to a handful. It is the right tool for that job and no worse than any alternative.
Stop using it the moment you are forecasting spend. At that point you have, or can measure, your own input and output token counts, and multiplying each by its own published rate is both easy and exact. A blended figure at the budgeting stage is an assumption standing in for data you already hold.
Frequently asked questions
What is blended cost?
A weighted average of a model's input and output prices, expressed as one figure per million tokens. Our default weights input at 70 per cent and output at 30 per cent. It is a calculated comparison metric rather than a price any provider publishes, and it exists so models with different input and output rates can be ranked on one axis.
Why is the default 70/30?
It approximates a common shape: a moderate prompt producing a longer answer. The specific split matters less than applying it consistently, since the figure is for ranking rather than forecasting. Workloads that depart from that shape should be priced on their own measured token mix.
Should I budget using a blended rate?
No. Price your measured input and output volumes separately against each published rate. Because output typically costs at least twice what input does, a blended figure can understate or overstate a real bill substantially depending on how much your workload generates.
See it in the data
Related guides
Terms used in this guide
Published by Tokenando. Last updated 2026-08-07. Figures in this guide are computed from our own pricing index and dated where they can move; see the methodology and corrections policy.