Claude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/MClaude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/M

Cheapest AI Models for Agents

Ranked from provider-published pricing · Prices checked 12 August 2026

Agentic workloads invert the RAG pattern: short instructions in, long chains of generated reasoning, tool calls and code out. Output tokens dominate the bill, so this list weights output at 80%.

Because providers typically charge several times more per output token than per input token, the cheapest agent models can look quite different from the cheapest models overall.

The ranking

top 20 of 271

Gemma 2 2B from Google leads this ranking at $0.020/M. That is 98% below the $0.900/M median across the 271 models that qualify for this list.

#ModelProviderOutput-Heavy WorkloadsInput /1MOutput /1MContext
1Gemma 2 2BGoogle$0.020/M$0.020$0.0208,000
2Qwen2.5-1.5B-InstructQwen$0.030/M$0.030$0.03032,000
3Llama 3.2 1BMeta$0.040/M$0.040$0.040128,000
4Phi-3-Mini-4K (NIM)Nvidia$0.040/M$0.040$0.0404,000
5Qwen2.5-3B-InstructQwen$0.050/M$0.050$0.05032,000
6ABAB 5.5cMiniMax$0.050/M$0.050$0.05016,000
7Llama 3.2 3BMeta$0.060/M$0.060$0.060128,000
8Llama 3.2 3B (Groq)Groq$0.060/M$0.060$0.060128,000
9Gemma 2 9B (DI)DeepInfra$0.060/M$0.060$0.0608,000
10Mistral 7B (DI)DeepInfra$0.070/M$0.070$0.07032,000
11Mistral 7B (Lepton)Lepton AI$0.070/M$0.070$0.07032,000
12Llama 3.1 8B (Groq)Groq$0.074/M$0.050$0.080128,000
13Doubao-Lite-128kByteDance$0.080/M$0.040$0.090128,000
14Doubao-Lite-32kByteDance$0.080/M$0.040$0.09032,000
15Mistral 7B v0.3Mistral$0.080/M$0.080$0.08032,000
16Granite 3.1 2B InstructIBM$0.086/M$0.030$0.100128,000
17Gemma 3 9BGoogle$0.090/M$0.090$0.090128,000
18Gemma 2 9BGoogle$0.090/M$0.090$0.0908,000
19Llama 3.1 8BMeta$0.100/M$0.100$0.100128,000
20Llama 3.1 8B (Together)Together AI$0.100/M$0.100$0.100128,000

How this list is built

  • Ranked cheapest first by a blended rate weighting input and output at 20/80. Output-heavy workloads such as long-form generation, agentic loops.
  • Excludes embed models.
  • Includes models marked available or preview; retired and deprecated SKUs are excluded.
  • Prices are the rates each provider publishes, not estimates. How the blended rate is calculated.

Frequently asked questions

Why do agents cost more to run?

An agent makes many model calls per task, and each one generates tokens: a plan, a tool call, a reflection on the result, then the next step. Output tokens are the expensive side of almost every price sheet, and an agent loop multiplies them, so a workload that looks affordable per call can become the dominant line item at scale.

Should I use a cheaper model for agent steps?

Often yes. A common pattern is to route planning and final synthesis to a strong model while handling routine intermediate steps (parsing, formatting, simple tool selection) with a cheap one. Since intermediate steps usually make up most of the calls, that mix captures most of the saving without changing the answer quality much.

Related rankings

Price it for your workload

Rankings are recomputed on every deploy from the catalogue. Prices checked 12 August 2026.