Cheapest Long-Context AI Models
Ranked from provider-published pricing · Prices checked 12 August 2026
Long context is what lets a model read a whole codebase, contract or research corpus in one request. This list keeps only models with a window of at least 200K tokens, then ranks them cheapest first.
Note that a big window and a cheap rate together can still be expensive in practice: filling a 200K-token window on every request costs 200K input tokens every time.
The ranking
top 20 of 70Gemini 1.5 Flash-8B from Google leads this ranking at $0.071/M. That is 98% below the $3.30/M median across the 70 models that qualify for this list.
| # | Model | Provider | 200K+ Context Ranked by Blended Cost | Input /1M | Output /1M | Context |
|---|---|---|---|---|---|---|
| 1 | Gemini 1.5 Flash-8B | $0.071/M | $0.037 | $0.150 | 1,000,000 | |
| 2 | Nova Lite | Amazon | $0.114/M | $0.060 | $0.240 | 300,000 |
| 3 | Gemini 2.0 Flash Lite | $0.142/M | $0.075 | $0.300 | 1,000,000 | |
| 4 | Gemini 1.5 Flash | $0.142/M | $0.075 | $0.300 | 1,000,000 | |
| 5 | Llama 4 Scout | Meta | $0.175/M | $0.100 | $0.350 | 512,000 |
| 6 | DeepSeek-V4-Flash | DeepSeek | $0.182/M | $0.140 | $0.280 | 1,000,000 |
| 7 | GPT-4.1 nano | OpenAI | $0.190/M | $0.100 | $0.400 | 1,000,000 |
| 8 | Gemini 2.0 Flash | $0.190/M | $0.100 | $0.400 | 1,000,000 | |
| 9 | Codestral Mamba | Mistral | $0.250/M | $0.250 | $0.250 | 256,000 |
| 10 | Jamba 1.5 Mini | AI21 Labs | $0.260/M | $0.200 | $0.400 | 256,000 |
| 11 | Gemini 2.5 Flash | $0.285/M | $0.150 | $0.600 | 1,000,000 | |
| 12 | Codestral | Mistral | $0.320/M | $0.200 | $0.600 | 256,000 |
| 13 | Hunyuan-Standard | Tencent | $0.455/M | $0.350 | $0.700 | 256,000 |
| 14 | GPT-5.6 Luna | OpenAI | $0.500/M | $0.200 | $1.20 | 1,050,000 |
| 15 | Jamba Instruct | AI21 Labs | $0.560/M | $0.500 | $0.700 | 256,000 |
| 16 | GLM-4-Long | Zhipu AI | $0.700/M | $0.700 | $0.700 | 1,000,000 |
| 17 | GPT-4.1 mini | OpenAI | $0.760/M | $0.400 | $1.60 | 1,000,000 |
| 18 | Qwen3.7-Plus | Qwen | $0.760/M | $0.400 | $1.60 | 1,000,000 |
| 19 | Mistral Large 3 | Mistral | $0.800/M | $0.500 | $1.50 | 256,000 |
| 20 | Doubao-Pro-256k | ByteDance | $0.812/M | $0.560 | $1.40 | 256,000 |
How this list is built
- Ranked cheapest first by a blended rate weighting input and output at 70/30. Default reference rate. 70% weight on input, 30% on output.
- Excludes embed models.
- Requires a context window of at least 200K tokens.
- Includes models marked available or preview; retired and deprecated SKUs are excluded.
- Prices are the rates each provider publishes, not estimates. How the blended rate is calculated.
Frequently asked questions
How much does a long-context request cost?
Multiply the tokens you actually send by the model’s input rate: filling a 200K-token window at $1.00 per million input tokens costs about $0.20 for that single request, before any output. Long context is priced per token used, not per window size, so sending less remains the cheapest optimisation.
Related rankings
Price it for your workload
Rankings are recomputed on every deploy from the catalogue. Prices checked 12 August 2026.