Cheapest Vision AI Models
Ranked from provider-published pricing · Prices checked 12 August 2026
These models accept images alongside text, for document understanding, screenshot analysis, chart reading and visual QA. Only models with published vision support are included, ranked cheapest first on blended token cost.
Image inputs are usually billed as tokens, with the count depending on resolution, so the token rates below drive image cost too, though the tokens-per-image conversion varies by provider.
The ranking
top 20 of 68Gemini 1.5 Flash-8B from Google leads this ranking at $0.071/M. That is 98% below the $3.40/M median across the 68 models that qualify for this list.
| # | Model | Provider | Image-Capable Models Ranked by Cost | Input /1M | Output /1M | Context |
|---|---|---|---|---|---|---|
| 1 | Gemini 1.5 Flash-8B | $0.071/M | $0.037 | $0.150 | 1,000,000 | |
| 2 | Gemma 3 9B | $0.090/M | $0.090 | $0.090 | 128,000 | |
| 3 | Nova Lite | Amazon | $0.114/M | $0.060 | $0.240 | 300,000 |
| 4 | Gemini 2.0 Flash Lite | $0.142/M | $0.075 | $0.300 | 1,000,000 | |
| 5 | Gemini 1.5 Flash | $0.142/M | $0.075 | $0.300 | 1,000,000 | |
| 6 | Pixtral 12B | Mistral | $0.150/M | $0.150 | $0.150 | 128,000 |
| 7 | Llama 3.2 11B Vision | Meta | $0.160/M | $0.160 | $0.160 | 128,000 |
| 8 | Llama 4 Scout | Meta | $0.175/M | $0.100 | $0.350 | 512,000 |
| 9 | Llama 3.2 11B Vision (Groq) | Groq | $0.180/M | $0.180 | $0.180 | 128,000 |
| 10 | GPT-4.1 nano | OpenAI | $0.190/M | $0.100 | $0.400 | 1,000,000 |
| 11 | Gemini 2.0 Flash | $0.190/M | $0.100 | $0.400 | 1,000,000 | |
| 12 | Gemma 3 27B | $0.270/M | $0.270 | $0.270 | 128,000 | |
| 13 | GPT-4o mini | OpenAI | $0.285/M | $0.150 | $0.600 | 128,000 |
| 14 | Gemini 2.5 Flash | $0.285/M | $0.150 | $0.600 | 1,000,000 | |
| 15 | Qwen2-VL-7B | Qwen | $0.300/M | $0.300 | $0.300 | 32,000 |
| 16 | Yi-VL-6B | 01.AI | $0.300/M | $0.300 | $0.300 | 4,000 |
| 17 | Llama 4 Maverick | Meta | $0.320/M | $0.200 | $0.600 | 128,000 |
| 18 | GPT-5.6 Luna | OpenAI | $0.500/M | $0.200 | $1.20 | 1,050,000 |
| 19 | GPT-4.1 mini | OpenAI | $0.760/M | $0.400 | $1.60 | 1,000,000 |
| 20 | Reka Edge | Reka AI | $0.880/M | $0.400 | $2.00 | 128,000 |
How this list is built
- Ranked cheapest first by a blended rate weighting input and output at 70/30. Default reference rate. 70% weight on input, 30% on output.
- Excludes embed models.
- Requires published vision support.
- Includes models marked available or preview; retired and deprecated SKUs are excluded.
- Prices are the rates each provider publishes, not estimates. How the blended rate is calculated.
Frequently asked questions
How are images priced in AI APIs?
Most providers convert an image into a number of input tokens based on its dimensions, then bill those at the standard input rate. A higher-resolution image becomes more tokens and costs more, which is why downscaling images before sending them is one of the most effective ways to cut a vision workload’s bill.
Related rankings
Price it for your workload
Rankings are recomputed on every deploy from the catalogue. Prices checked 12 August 2026.