Learn · AI economics
AI Pricing Guides
16 guidesHow AI API pricing and model measurement actually work, written for the people who pay the bill. For single-term definitions see the AI glossary; for live rates, the pricing index.
Guides
16How AI API Pricing WorksAI APIs bill per token, charge more for output than input, and range over 1,600× in price. Here is what actually drives the number on your invoice.Input vs Output Tokens: Why the Split MattersOutput tokens cost 2× more than input on the median AI model, and 3× or more on 117 of them. Why that asymmetry decides which model is cheapest for you.Blended Cost ExplainedA blended rate collapses input and output prices into one number for ranking. Useful for shortlisting, misleading for budgeting, and here is why.Cost per 1,000 Tokens vs per MillionAI prices moved from per-1,000 to per-million tokens. The conversion is a factor of 1,000, and mixing the two units is a common budgeting error.How to Calculate AI API CostsThe arithmetic for estimating LLM API spend, a worked example, and the four mistakes that make most estimates too low.How Context Windows Affect CostA large context window is a capacity, not a bill. What costs money is filling it, and several providers charge a higher rate past a threshold.Prompt Caching: When It Saves 90%Cached input typically costs 10% of the standard rate, but cache writes cost more than a normal request. The break-even maths and when caching pays.Batch API Pricing ExplainedBatch processing halves both input and output rates in exchange for asynchronous delivery. No quality trade-off, which makes it unusually easy to justify.Free AI APIs vs Paid: What You Give UpGenuinely free AI models exist. The constraints are rate limits, availability and support rather than output quality, which changes where they fit.Open-Weights vs API: The Real Cost ComparisonSelf-hosting an open-weights model trades a per-token bill for a GPU bill. Where the crossover sits depends almost entirely on utilisation.Estimating RAG CostsRetrieval-augmented generation is input-heavy, so the input rate dominates. Embedding is the small line; resent context is the large one.What Agentic Workloads Actually CostAgents multiply output tokens, the expensive side of every price sheet, and reasoning models bill for thinking you never see.How to Reduce LLM API CostsThe levers that actually cut AI API spend, ranked by impact, starting with the one that dwarfs all the others.Are AI Prices Actually Falling?We have logged every tracked price change since May. Cuts outnumber rises, but the average move is close to flat: prices churn rather than fall.How to Choose an AI ModelA decision order that works: rule out on capacity, shortlist on price, break the tie on evidence, then verify on your own traffic.Why AI Benchmark Scores Do not CompareTwo AI models can post benchmark scores that look comparable and are not. Here is what breaks comparability, and how to read scores safely.