Claude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/MClaude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/M

Cheapest AI Models for Chat

Ranked from provider-published pricing · Prices checked 12 August 2026

Conversational workloads sit between the input-heavy and output-heavy extremes: a growing transcript goes in, a moderate reply comes out. This list ranks on an even 50/50 weighting for that balance.

The ranking

top 20 of 271

Gemma 2 2B from Google leads this ranking at $0.020/M. That is 98% below the $0.800/M median across the 271 models that qualify for this list.

#ModelProvidera Balanced 50/50 BlendInput /1MOutput /1MContext
1Gemma 2 2BGoogle$0.020/M$0.020$0.0208,000
2Qwen2.5-1.5B-InstructQwen$0.030/M$0.030$0.03032,000
3Llama 3.2 1BMeta$0.040/M$0.040$0.040128,000
4Phi-3-Mini-4K (NIM)Nvidia$0.040/M$0.040$0.0404,000
5Qwen2.5-3B-InstructQwen$0.050/M$0.050$0.05032,000
6ABAB 5.5cMiniMax$0.050/M$0.050$0.05016,000
7Llama 3.2 3BMeta$0.060/M$0.060$0.060128,000
8Llama 3.2 3B (Groq)Groq$0.060/M$0.060$0.060128,000
9Gemma 2 9B (DI)DeepInfra$0.060/M$0.060$0.0608,000
10Llama 3.1 8B (Groq)Groq$0.065/M$0.050$0.080128,000
11Granite 3.1 2B InstructIBM$0.065/M$0.030$0.100128,000
12Doubao-Lite-128kByteDance$0.065/M$0.040$0.090128,000
13Doubao-Lite-32kByteDance$0.065/M$0.040$0.09032,000
14Mistral 7B (DI)DeepInfra$0.070/M$0.070$0.07032,000
15Mistral 7B (Lepton)Lepton AI$0.070/M$0.070$0.07032,000
16Mistral 7B v0.3Mistral$0.080/M$0.080$0.08032,000
17Nova MicroAmazon$0.088/M$0.035$0.140128,000
18Gemma 3 9BGoogle$0.090/M$0.090$0.090128,000
19Gemma 2 9BGoogle$0.090/M$0.090$0.0908,000
20Gemini 1.5 Flash-8BGoogle$0.094/M$0.037$0.1501,000,000

How this list is built

  • Ranked cheapest first by a blended rate weighting input and output at 50/50. Symmetrical blend. Useful for back-and-forth chat workloads.
  • Excludes embed models.
  • Includes models marked available or preview; retired and deprecated SKUs are excluded.
  • Prices are the rates each provider publishes, not estimates. How the blended rate is calculated.

Frequently asked questions

What is the cheapest model for a chatbot?

On an even weighting of input and output token prices, the models at the top of this table are the cheapest we track. Remember that a chat transcript grows with every turn, so the input side of a long conversation costs more than it does on the first message, a large context window and cached-input pricing both help.

Related rankings

Price it for your workload

Rankings are recomputed on every deploy from the catalogue. Prices checked 12 August 2026.