Claude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/MClaude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/M

Cheapest AI Models

Ranked from provider-published pricing · Prices checked 12 August 2026

These are the cheapest AI models we track, ranked by blended cost per million tokens, a single figure that weights a model’s input rate at 70% and its output rate at 30%, so models with very different input/output splits can be compared on one axis.

Cheapest overall is not the same as cheapest for your workload. A model with a low input rate and an expensive output rate wins here but loses badly on long-form generation. If you know your token mix, the input-heavy and output-heavy lists below rank the same catalogue on those weightings instead.

The ranking

top 20 of 271

Gemma 2 2B from Google leads this ranking at $0.020/M. That is 98% below the $0.800/M median across the 271 models that qualify for this list.

#ModelProviderBlended Cost per 1M TokensInput /1MOutput /1MContext
1Gemma 2 2BGoogle$0.020/M$0.020$0.0208,000
2Qwen2.5-1.5B-InstructQwen$0.030/M$0.030$0.03032,000
3Llama 3.2 1BMeta$0.040/M$0.040$0.040128,000
4Phi-3-Mini-4K (NIM)Nvidia$0.040/M$0.040$0.0404,000
5Qwen2.5-3B-InstructQwen$0.050/M$0.050$0.05032,000
6ABAB 5.5cMiniMax$0.050/M$0.050$0.05016,000
7Granite 3.1 2B InstructIBM$0.051/M$0.030$0.100128,000
8Doubao-Lite-128kByteDance$0.055/M$0.040$0.090128,000
9Doubao-Lite-32kByteDance$0.055/M$0.040$0.09032,000
10Llama 3.1 8B (Groq)Groq$0.059/M$0.050$0.080128,000
11Llama 3.2 3BMeta$0.060/M$0.060$0.060128,000
12Llama 3.2 3B (Groq)Groq$0.060/M$0.060$0.060128,000
13Gemma 2 9B (DI)DeepInfra$0.060/M$0.060$0.0608,000
14Nova MicroAmazon$0.067/M$0.035$0.140128,000
15Mistral 7B (DI)DeepInfra$0.070/M$0.070$0.07032,000
16Mistral 7B (Lepton)Lepton AI$0.070/M$0.070$0.07032,000
17Gemini 1.5 Flash-8BGoogle$0.071/M$0.037$0.1501,000,000
18Command R7BCohere$0.071/M$0.037$0.150128,000
19Phi-4 MiniMicrosoft$0.076/M$0.040$0.160128,000
20Mistral 7B v0.3Mistral$0.080/M$0.080$0.08032,000

How this list is built

  • Ranked cheapest first by a blended rate weighting input and output at 70/30. Default reference rate. 70% weight on input, 30% on output.
  • Excludes embed models.
  • Includes models marked available or preview; retired and deprecated SKUs are excluded.
  • Prices are the rates each provider publishes, not estimates. How the blended rate is calculated.

Frequently asked questions

What is the cheapest AI model?

By blended cost per million tokens (70% input, 30% output) the cheapest model in our catalogue is the one ranked first in the table above. Because providers change prices frequently and free tiers come and go, treat the top few as a shortlist rather than a permanent winner, and check the per-model page for the current published rate and its source.

Why rank by blended cost instead of input price?

Input price alone is easy to game: a provider can advertise a very low input rate while charging many times more per output token. Because almost every real workload pays for both, a blended figure is a fairer single-number comparison. It is still an assumption, not a measurement, price your own token volumes for an accurate answer.

Are cheap AI models worse?

Not automatically. Cheap models are usually smaller or more heavily optimised, which shows up on hard reasoning and coding benchmarks, but for classification, extraction, summarisation and routing they are frequently indistinguishable from frontier models at a fraction of the cost. The benchmark-led lists below show where the cheap end still scores well.

Related rankings

Price it for your workload

Rankings are recomputed on every deploy from the catalogue. Prices checked 12 August 2026.