Cheapest AI Models for Chat

Ranked from provider-published pricing · Prices checked 27 September 2026

Conversational workloads sit between the input-heavy and output-heavy extremes: a growing transcript goes in, a moderate reply comes out. This list ranks on an even 50/50 weighting for that balance.

The ranking

top 20 of 271

Gemma 2 2B from Google leads this ranking at $0.020/M. That is 98% below the $0.800/M median across the 271 models that qualify for this list.

#ModelProvidera Balanced 50/50 BlendInput /1MOutput /1MContext
1Gemma 2 2BGoogle$0.020/M$0.020$0.0208,000
2Qwen2.5-1.5B-InstructQwen$0.030/M$0.030$0.03032,000
3Llama 3.2 1BMeta$0.040/M$0.040$0.040128,000
4Phi-3-Mini-4K (NIM)Nvidia$0.040/M$0.040$0.0404,000
5Qwen2.5-3B-InstructQwen$0.050/M$0.050$0.05032,000
6ABAB 5.5cMiniMax$0.050/M$0.050$0.05016,000
7Llama 3.2 3BMeta$0.060/M$0.060$0.060128,000
8Llama 3.2 3B (Groq)Groq$0.060/M$0.060$0.060128,000
9Gemma 2 9B (DI)DeepInfra$0.060/M$0.060$0.0608,000
10Llama 3.1 8B (Groq)Groq$0.065/M$0.050$0.080128,000
11Granite 3.1 2B InstructIBM$0.065/M$0.030$0.100128,000
12Doubao-Lite-128kByteDance$0.065/M$0.040$0.090128,000
13Doubao-Lite-32kByteDance$0.065/M$0.040$0.09032,000
14Mistral 7B (DI)DeepInfra$0.070/M$0.070$0.07032,000
15Mistral 7B (Lepton)Lepton AI$0.070/M$0.070$0.07032,000
16Mistral 7B v0.3Mistral$0.080/M$0.080$0.08032,000
17Nova MicroAmazon$0.088/M$0.035$0.140128,000
18Gemma 3 9BGoogle$0.090/M$0.090$0.090128,000
19Gemma 2 9BGoogle$0.090/M$0.090$0.0908,000
20Gemini 1.5 Flash-8BGoogle$0.094/M$0.037$0.1501,000,000

How this list is built

  • Ranked cheapest first by a blended rate weighting input and output at 50/50. Symmetrical blend. Useful for back-and-forth chat workloads.
  • Excludes embed models.
  • Includes models marked available or preview; retired and deprecated SKUs are excluded.
  • Prices are the rates each provider publishes, not estimates. How the blended rate is calculated.

Frequently asked questions

What is the cheapest model for a chatbot?

On an even weighting of input and output token prices, the models at the top of this table are the cheapest we track. Remember that a chat transcript grows with every turn, so the input side of a long conversation costs more than it does on the first message, a large context window and cached-input pricing both help.

Related rankings

Price it for your workload

Rankings are recomputed on every deploy from the catalogue. Prices checked 27 September 2026.