Cheapest AI Models for Agents
Ranked from provider-published pricing · Prices checked 12 August 2026
Agentic workloads invert the RAG pattern: short instructions in, long chains of generated reasoning, tool calls and code out. Output tokens dominate the bill, so this list weights output at 80%.
Because providers typically charge several times more per output token than per input token, the cheapest agent models can look quite different from the cheapest models overall.
The ranking
top 20 of 271Gemma 2 2B from Google leads this ranking at $0.020/M. That is 98% below the $0.900/M median across the 271 models that qualify for this list.
| # | Model | Provider | Output-Heavy Workloads | Input /1M | Output /1M | Context |
|---|---|---|---|---|---|---|
| 1 | Gemma 2 2B | $0.020/M | $0.020 | $0.020 | 8,000 | |
| 2 | Qwen2.5-1.5B-Instruct | Qwen | $0.030/M | $0.030 | $0.030 | 32,000 |
| 3 | Llama 3.2 1B | Meta | $0.040/M | $0.040 | $0.040 | 128,000 |
| 4 | Phi-3-Mini-4K (NIM) | Nvidia | $0.040/M | $0.040 | $0.040 | 4,000 |
| 5 | Qwen2.5-3B-Instruct | Qwen | $0.050/M | $0.050 | $0.050 | 32,000 |
| 6 | ABAB 5.5c | MiniMax | $0.050/M | $0.050 | $0.050 | 16,000 |
| 7 | Llama 3.2 3B | Meta | $0.060/M | $0.060 | $0.060 | 128,000 |
| 8 | Llama 3.2 3B (Groq) | Groq | $0.060/M | $0.060 | $0.060 | 128,000 |
| 9 | Gemma 2 9B (DI) | DeepInfra | $0.060/M | $0.060 | $0.060 | 8,000 |
| 10 | Mistral 7B (DI) | DeepInfra | $0.070/M | $0.070 | $0.070 | 32,000 |
| 11 | Mistral 7B (Lepton) | Lepton AI | $0.070/M | $0.070 | $0.070 | 32,000 |
| 12 | Llama 3.1 8B (Groq) | Groq | $0.074/M | $0.050 | $0.080 | 128,000 |
| 13 | Doubao-Lite-128k | ByteDance | $0.080/M | $0.040 | $0.090 | 128,000 |
| 14 | Doubao-Lite-32k | ByteDance | $0.080/M | $0.040 | $0.090 | 32,000 |
| 15 | Mistral 7B v0.3 | Mistral | $0.080/M | $0.080 | $0.080 | 32,000 |
| 16 | Granite 3.1 2B Instruct | IBM | $0.086/M | $0.030 | $0.100 | 128,000 |
| 17 | Gemma 3 9B | $0.090/M | $0.090 | $0.090 | 128,000 | |
| 18 | Gemma 2 9B | $0.090/M | $0.090 | $0.090 | 8,000 | |
| 19 | Llama 3.1 8B | Meta | $0.100/M | $0.100 | $0.100 | 128,000 |
| 20 | Llama 3.1 8B (Together) | Together AI | $0.100/M | $0.100 | $0.100 | 128,000 |
How this list is built
- Ranked cheapest first by a blended rate weighting input and output at 20/80. Output-heavy workloads such as long-form generation, agentic loops.
- Excludes embed models.
- Includes models marked available or preview; retired and deprecated SKUs are excluded.
- Prices are the rates each provider publishes, not estimates. How the blended rate is calculated.
Frequently asked questions
Why do agents cost more to run?
An agent makes many model calls per task, and each one generates tokens: a plan, a tool call, a reflection on the result, then the next step. Output tokens are the expensive side of almost every price sheet, and an agent loop multiplies them, so a workload that looks affordable per call can become the dominant line item at scale.
Should I use a cheaper model for agent steps?
Often yes. A common pattern is to route planning and final synthesis to a strong model while handling routine intermediate steps (parsing, formatting, simple tool selection) with a cheap one. Since intermediate steps usually make up most of the calls, that mix captures most of the saving without changing the answer quality much.
Related rankings
Price it for your workload
Rankings are recomputed on every deploy from the catalogue. Prices checked 12 August 2026.