Cheapest AI Models
Ranked from provider-published pricing · Prices checked 12 August 2026
These are the cheapest AI models we track, ranked by blended cost per million tokens, a single figure that weights a model’s input rate at 70% and its output rate at 30%, so models with very different input/output splits can be compared on one axis.
Cheapest overall is not the same as cheapest for your workload. A model with a low input rate and an expensive output rate wins here but loses badly on long-form generation. If you know your token mix, the input-heavy and output-heavy lists below rank the same catalogue on those weightings instead.
The ranking
top 20 of 271Gemma 2 2B from Google leads this ranking at $0.020/M. That is 98% below the $0.800/M median across the 271 models that qualify for this list.
| # | Model | Provider | Blended Cost per 1M Tokens | Input /1M | Output /1M | Context |
|---|---|---|---|---|---|---|
| 1 | Gemma 2 2B | $0.020/M | $0.020 | $0.020 | 8,000 | |
| 2 | Qwen2.5-1.5B-Instruct | Qwen | $0.030/M | $0.030 | $0.030 | 32,000 |
| 3 | Llama 3.2 1B | Meta | $0.040/M | $0.040 | $0.040 | 128,000 |
| 4 | Phi-3-Mini-4K (NIM) | Nvidia | $0.040/M | $0.040 | $0.040 | 4,000 |
| 5 | Qwen2.5-3B-Instruct | Qwen | $0.050/M | $0.050 | $0.050 | 32,000 |
| 6 | ABAB 5.5c | MiniMax | $0.050/M | $0.050 | $0.050 | 16,000 |
| 7 | Granite 3.1 2B Instruct | IBM | $0.051/M | $0.030 | $0.100 | 128,000 |
| 8 | Doubao-Lite-128k | ByteDance | $0.055/M | $0.040 | $0.090 | 128,000 |
| 9 | Doubao-Lite-32k | ByteDance | $0.055/M | $0.040 | $0.090 | 32,000 |
| 10 | Llama 3.1 8B (Groq) | Groq | $0.059/M | $0.050 | $0.080 | 128,000 |
| 11 | Llama 3.2 3B | Meta | $0.060/M | $0.060 | $0.060 | 128,000 |
| 12 | Llama 3.2 3B (Groq) | Groq | $0.060/M | $0.060 | $0.060 | 128,000 |
| 13 | Gemma 2 9B (DI) | DeepInfra | $0.060/M | $0.060 | $0.060 | 8,000 |
| 14 | Nova Micro | Amazon | $0.067/M | $0.035 | $0.140 | 128,000 |
| 15 | Mistral 7B (DI) | DeepInfra | $0.070/M | $0.070 | $0.070 | 32,000 |
| 16 | Mistral 7B (Lepton) | Lepton AI | $0.070/M | $0.070 | $0.070 | 32,000 |
| 17 | Gemini 1.5 Flash-8B | $0.071/M | $0.037 | $0.150 | 1,000,000 | |
| 18 | Command R7B | Cohere | $0.071/M | $0.037 | $0.150 | 128,000 |
| 19 | Phi-4 Mini | Microsoft | $0.076/M | $0.040 | $0.160 | 128,000 |
| 20 | Mistral 7B v0.3 | Mistral | $0.080/M | $0.080 | $0.080 | 32,000 |
How this list is built
- Ranked cheapest first by a blended rate weighting input and output at 70/30. Default reference rate. 70% weight on input, 30% on output.
- Excludes embed models.
- Includes models marked available or preview; retired and deprecated SKUs are excluded.
- Prices are the rates each provider publishes, not estimates. How the blended rate is calculated.
Frequently asked questions
What is the cheapest AI model?
By blended cost per million tokens (70% input, 30% output) the cheapest model in our catalogue is the one ranked first in the table above. Because providers change prices frequently and free tiers come and go, treat the top few as a shortlist rather than a permanent winner, and check the per-model page for the current published rate and its source.
Why rank by blended cost instead of input price?
Input price alone is easy to game: a provider can advertise a very low input rate while charging many times more per output token. Because almost every real workload pays for both, a blended figure is a fairer single-number comparison. It is still an assumption, not a measurement, price your own token volumes for an accurate answer.
Are cheap AI models worse?
Not automatically. Cheap models are usually smaller or more heavily optimised, which shows up on hard reasoning and coding benchmarks, but for classification, extraction, summarisation and routing they are frequently indistinguishable from frontier models at a fraction of the cost. The benchmark-led lists below show where the cheap end still scores well.
Related rankings
Price it for your workload
Rankings are recomputed on every deploy from the catalogue. Prices checked 12 August 2026.