Cheapest AI Models for Chat
Ranked from provider-published pricing · Prices checked 12 August 2026
Conversational workloads sit between the input-heavy and output-heavy extremes: a growing transcript goes in, a moderate reply comes out. This list ranks on an even 50/50 weighting for that balance.
The ranking
top 20 of 271Gemma 2 2B from Google leads this ranking at $0.020/M. That is 98% below the $0.800/M median across the 271 models that qualify for this list.
| # | Model | Provider | a Balanced 50/50 Blend | Input /1M | Output /1M | Context |
|---|---|---|---|---|---|---|
| 1 | Gemma 2 2B | $0.020/M | $0.020 | $0.020 | 8,000 | |
| 2 | Qwen2.5-1.5B-Instruct | Qwen | $0.030/M | $0.030 | $0.030 | 32,000 |
| 3 | Llama 3.2 1B | Meta | $0.040/M | $0.040 | $0.040 | 128,000 |
| 4 | Phi-3-Mini-4K (NIM) | Nvidia | $0.040/M | $0.040 | $0.040 | 4,000 |
| 5 | Qwen2.5-3B-Instruct | Qwen | $0.050/M | $0.050 | $0.050 | 32,000 |
| 6 | ABAB 5.5c | MiniMax | $0.050/M | $0.050 | $0.050 | 16,000 |
| 7 | Llama 3.2 3B | Meta | $0.060/M | $0.060 | $0.060 | 128,000 |
| 8 | Llama 3.2 3B (Groq) | Groq | $0.060/M | $0.060 | $0.060 | 128,000 |
| 9 | Gemma 2 9B (DI) | DeepInfra | $0.060/M | $0.060 | $0.060 | 8,000 |
| 10 | Llama 3.1 8B (Groq) | Groq | $0.065/M | $0.050 | $0.080 | 128,000 |
| 11 | Granite 3.1 2B Instruct | IBM | $0.065/M | $0.030 | $0.100 | 128,000 |
| 12 | Doubao-Lite-128k | ByteDance | $0.065/M | $0.040 | $0.090 | 128,000 |
| 13 | Doubao-Lite-32k | ByteDance | $0.065/M | $0.040 | $0.090 | 32,000 |
| 14 | Mistral 7B (DI) | DeepInfra | $0.070/M | $0.070 | $0.070 | 32,000 |
| 15 | Mistral 7B (Lepton) | Lepton AI | $0.070/M | $0.070 | $0.070 | 32,000 |
| 16 | Mistral 7B v0.3 | Mistral | $0.080/M | $0.080 | $0.080 | 32,000 |
| 17 | Nova Micro | Amazon | $0.088/M | $0.035 | $0.140 | 128,000 |
| 18 | Gemma 3 9B | $0.090/M | $0.090 | $0.090 | 128,000 | |
| 19 | Gemma 2 9B | $0.090/M | $0.090 | $0.090 | 8,000 | |
| 20 | Gemini 1.5 Flash-8B | $0.094/M | $0.037 | $0.150 | 1,000,000 |
How this list is built
- Ranked cheapest first by a blended rate weighting input and output at 50/50. Symmetrical blend. Useful for back-and-forth chat workloads.
- Excludes embed models.
- Includes models marked available or preview; retired and deprecated SKUs are excluded.
- Prices are the rates each provider publishes, not estimates. How the blended rate is calculated.
Frequently asked questions
What is the cheapest model for a chatbot?
On an even weighting of input and output token prices, the models at the top of this table are the cheapest we track. Remember that a chat transcript grows with every turn, so the input side of a long conversation costs more than it does on the first message, a large context window and cached-input pricing both help.
Related rankings
Price it for your workload
Rankings are recomputed on every deploy from the catalogue. Prices checked 12 August 2026.