Claude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/MClaude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/M

Free AI Models

Ranked from provider-published pricing · Prices checked 12 August 2026

These models are published at zero cost for both input and output tokens. That usually means an open-weights model hosted on a free tier, and free tiers normally carry rate limits, smaller context allowances or no availability guarantee.

A model only appears here when both its input and output prices are published as $0. Where a price is simply unknown we leave the model out rather than guess it is free.

The ranking

top 16 of 16

Nemotron 3.5 Lightning (free) from Nvidia leads this ranking at 1,000,000 tokens. The median across the 16 qualifying models is 259,000 tokens.

#ModelProviderModels Published at $0 per Million TokensInput /1MOutput /1M
1Nemotron 3.5 Lightning (free)Nvidia1,000,000 tokens$0$0
2Nemotron 3 Ultra (free)Nvidia1,000,000 tokens$0$0
3Ling 3.0 Tiny (free)inclusionai262,000 tokens$0$0
4Laguna S 2.1 (free)poolside262,000 tokens$0$0
5Laguna XS 2.1 (free)poolside262,000 tokens$0$0
6Gemma 4 26B A4B (free)Google262,000 tokens$0$0
7Gemma 4 31B (free)Google262,000 tokens$0$0
8Nemotron 3 Super (free)Nvidia262,000 tokens$0$0
9North Mini Code (free)Cohere256,000 tokens$0$0
10Nemotron 3 Nano Omni (free)Nvidia256,000 tokens$0$0
11Nemotron 3 Nano 30B A3B (free)Nvidia256,000 tokens$0$0
12gpt-oss-20b (free)OpenAI131,000 tokens$0$0
13LiquidAI: LFM2.5-2.6B (free)liquid128,000 tokens$0$0
14Nemotron 3.5 Content Safety (free)Nvidia128,000 tokens$0$0
15Nemotron Nano 12B 2 VL (free)Nvidia128,000 tokens$0$0
16Nemotron Nano 9B V2 (free)Nvidia128,000 tokens$0$0

How this list is built

  • Ranked by maximum context window, largest first.
  • Only models whose published input and output prices are both $0.
  • Includes models marked available or preview; retired and deprecated SKUs are excluded.
  • Prices are the rates each provider publishes, not estimates. How the blended rate is calculated.

Frequently asked questions

Are there genuinely free AI APIs?

Yes, though almost always with limits. Open-weights models are frequently offered at no per-token charge on a hosted free tier, subject to rate limits and without an availability guarantee. For production workloads treat a free tier as a way to prototype, and price the paid tier before you depend on it.

Is a free model good enough for production?

For well-scoped tasks (classification, extraction, routing, simple summarisation) open-weights models are often entirely sufficient. The tradeoffs are usually operational rather than qualitative: rate limits, latency variability, and no contractual support if the endpoint goes away.

Related rankings

Price it for your workload

Rankings are recomputed on every deploy from the catalogue. Prices checked 12 August 2026.