Cheapest AI Models for RAG
Ranked from provider-published pricing · Prices checked 12 August 2026
Retrieval-augmented generation is input-heavy: you send large retrieved passages and get back a short answer. Ranking on the standard blend understates that, so this list weights input at 80% and output at 20%.
The ordering changes meaningfully against the headline cheapest list, models with cheap input and expensive output rise here, which is exactly the shape a RAG workload wants.
The ranking
top 20 of 271Gemma 2 2B from Google leads this ranking at $0.020/M. That is 97% below the $0.700/M median across the 271 models that qualify for this list.
| # | Model | Provider | Input-Heavy Workloads | Input /1M | Output /1M | Context |
|---|---|---|---|---|---|---|
| 1 | Gemma 2 2B | $0.020/M | $0.020 | $0.020 | 8,000 | |
| 2 | Qwen2.5-1.5B-Instruct | Qwen | $0.030/M | $0.030 | $0.030 | 32,000 |
| 3 | Llama 3.2 1B | Meta | $0.040/M | $0.040 | $0.040 | 128,000 |
| 4 | Phi-3-Mini-4K (NIM) | Nvidia | $0.040/M | $0.040 | $0.040 | 4,000 |
| 5 | Granite 3.1 2B Instruct | IBM | $0.044/M | $0.030 | $0.100 | 128,000 |
| 6 | Doubao-Lite-128k | ByteDance | $0.050/M | $0.040 | $0.090 | 128,000 |
| 7 | Doubao-Lite-32k | ByteDance | $0.050/M | $0.040 | $0.090 | 32,000 |
| 8 | Qwen2.5-3B-Instruct | Qwen | $0.050/M | $0.050 | $0.050 | 32,000 |
| 9 | ABAB 5.5c | MiniMax | $0.050/M | $0.050 | $0.050 | 16,000 |
| 10 | Nova Micro | Amazon | $0.056/M | $0.035 | $0.140 | 128,000 |
| 11 | Llama 3.1 8B (Groq) | Groq | $0.056/M | $0.050 | $0.080 | 128,000 |
| 12 | Gemini 1.5 Flash-8B | $0.060/M | $0.037 | $0.150 | 1,000,000 | |
| 13 | Llama 3.2 3B | Meta | $0.060/M | $0.060 | $0.060 | 128,000 |
| 14 | Command R7B | Cohere | $0.060/M | $0.037 | $0.150 | 128,000 |
| 15 | Llama 3.2 3B (Groq) | Groq | $0.060/M | $0.060 | $0.060 | 128,000 |
| 16 | Gemma 2 9B (DI) | DeepInfra | $0.060/M | $0.060 | $0.060 | 8,000 |
| 17 | Phi-4 Mini | Microsoft | $0.064/M | $0.040 | $0.160 | 128,000 |
| 18 | Mistral 7B (DI) | DeepInfra | $0.070/M | $0.070 | $0.070 | 32,000 |
| 19 | Mistral 7B (Lepton) | Lepton AI | $0.070/M | $0.070 | $0.070 | 32,000 |
| 20 | Mistral 7B v0.3 | Mistral | $0.080/M | $0.080 | $0.080 | 32,000 |
How this list is built
- Ranked cheapest first by a blended rate weighting input and output at 80/20. Input-heavy workloads such as long-context RAG or document Q&A.
- Excludes embed models.
- Includes models marked available or preview; retired and deprecated SKUs are excluded.
- Prices are the rates each provider publishes, not estimates. How the blended rate is calculated.
Frequently asked questions
What makes a model cheap for RAG?
A low input price, because RAG sends far more tokens than it receives: retrieved context, system instructions and the question go in, and a short grounded answer comes out. A large context window matters too, since it caps how many retrieved passages you can supply in one request.
Does prompt caching change these rankings?
Yes, often substantially. If your retrieved context repeats across requests, providers offering cached-input pricing can cut the dominant cost of a RAG workload by a large margin. This list ranks on standard published input rates; check the individual model pages for cached-input rates where the provider publishes them.
Related rankings
Price it for your workload
Rankings are recomputed on every deploy from the catalogue. Prices checked 12 August 2026.