Claude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/MClaude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/M

Cheapest Long-Context AI Models

Ranked from provider-published pricing · Prices checked 12 August 2026

Long context is what lets a model read a whole codebase, contract or research corpus in one request. This list keeps only models with a window of at least 200K tokens, then ranks them cheapest first.

Note that a big window and a cheap rate together can still be expensive in practice: filling a 200K-token window on every request costs 200K input tokens every time.

The ranking

top 20 of 70

Gemini 1.5 Flash-8B from Google leads this ranking at $0.071/M. That is 98% below the $3.30/M median across the 70 models that qualify for this list.

#ModelProvider200K+ Context Ranked by Blended CostInput /1MOutput /1MContext
1Gemini 1.5 Flash-8BGoogle$0.071/M$0.037$0.1501,000,000
2Nova LiteAmazon$0.114/M$0.060$0.240300,000
3Gemini 2.0 Flash LiteGoogle$0.142/M$0.075$0.3001,000,000
4Gemini 1.5 FlashGoogle$0.142/M$0.075$0.3001,000,000
5Llama 4 ScoutMeta$0.175/M$0.100$0.350512,000
6DeepSeek-V4-FlashDeepSeek$0.182/M$0.140$0.2801,000,000
7GPT-4.1 nanoOpenAI$0.190/M$0.100$0.4001,000,000
8Gemini 2.0 FlashGoogle$0.190/M$0.100$0.4001,000,000
9Codestral MambaMistral$0.250/M$0.250$0.250256,000
10Jamba 1.5 MiniAI21 Labs$0.260/M$0.200$0.400256,000
11Gemini 2.5 FlashGoogle$0.285/M$0.150$0.6001,000,000
12CodestralMistral$0.320/M$0.200$0.600256,000
13Hunyuan-StandardTencent$0.455/M$0.350$0.700256,000
14GPT-5.6 LunaOpenAI$0.500/M$0.200$1.201,050,000
15Jamba InstructAI21 Labs$0.560/M$0.500$0.700256,000
16GLM-4-LongZhipu AI$0.700/M$0.700$0.7001,000,000
17GPT-4.1 miniOpenAI$0.760/M$0.400$1.601,000,000
18Qwen3.7-PlusQwen$0.760/M$0.400$1.601,000,000
19Mistral Large 3Mistral$0.800/M$0.500$1.50256,000
20Doubao-Pro-256kByteDance$0.812/M$0.560$1.40256,000

How this list is built

  • Ranked cheapest first by a blended rate weighting input and output at 70/30. Default reference rate. 70% weight on input, 30% on output.
  • Excludes embed models.
  • Requires a context window of at least 200K tokens.
  • Includes models marked available or preview; retired and deprecated SKUs are excluded.
  • Prices are the rates each provider publishes, not estimates. How the blended rate is calculated.

Frequently asked questions

How much does a long-context request cost?

Multiply the tokens you actually send by the model’s input rate: filling a 200K-token window at $1.00 per million input tokens costs about $0.20 for that single request, before any output. Long context is priced per token used, not per window size, so sending less remains the cheapest optimisation.

Related rankings

Price it for your workload

Rankings are recomputed on every deploy from the catalogue. Prices checked 12 August 2026.