Claude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/MClaude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/M

Best Small AI Models

Ranked from provider-published pricing · Prices checked 12 August 2026

Small models trade some peak capability for much lower cost and latency. This list ranks the efficient tier by LMArena Elo, a crowd-sourced rating from blind head-to-head comparisons, so the ordering reflects human preference rather than a single test.

The best of these land close to frontier models on everyday tasks while costing a fraction as much, which is what makes them the default choice for high-volume production work.

The ranking

top 20 of 69

GPT-5.6 Luna from OpenAI leads this ranking at 1451 Elo. The median across the 69 qualifying models is 1184 Elo.

#ModelProviderEfficient Models Ranked by LMArena EloMMLUInput /1MOutput /1MBlended 70/30
1GPT-5.6 LunaOpenAI1451 Elo$0.200$1.20$0.500
2GPT-4.1 miniOpenAI1290 Elo81$0.400$1.60$0.760
3Claude Haiku 4.5Anthropic1280 Elo78$0.800$4.00$1.76
4GPT-4o miniOpenAI1273 Elo82$0.150$0.600$0.285
5Gemini 2.0 FlashGoogle1270 Elo78$0.100$0.400$0.190
6Grok 3 minixAI1265 Elo79$0.300$0.500$0.360
7Llama 4 ScoutMeta1265 Elo80$0.100$0.350$0.175
8Llama 3.3 70BMeta1257 Elo86$0.230$0.400$0.281
9Llama 3.3 70B (Groq)Groq1257 Elo86$0.590$0.790$0.650
10Llama 3.3 70B (Together)Together AI1257 Elo86$0.880$0.880$0.880
11Llama 3.3 70B (DI)DeepInfra1257 Elo86$0.230$0.400$0.281
12Llama-3.3-70B (Cerebras)Cerebras1257 Elo86$0.590$0.990$0.710
13Llama 3.1 70BMeta1248 Elo83$0.350$0.400$0.365
14Llama 3.1 70B (Groq)Groq1248 Elo83$0.590$0.790$0.650
15Llama 3.1 70B (DI)DeepInfra1248 Elo83$0.350$0.400$0.365
16Llama 3.1 70B (Hyp)Hyperbolic1248 Elo83$0.400$0.400$0.400
17Llama 3.1 70B (Anyscale)Anyscale1248 Elo83$1.00$1.00$1.00
18Llama 3.1 70B (Lepton)Lepton AI1248 Elo83$0.800$0.800$0.800
19Llama 3.1 70B (DB)Databricks1248 Elo83$1.00$3.00$1.60
20Yi-Lightning01.AI1244 Elo80$0.140$0.140$0.140

How this list is built

  • Ranked by LMArena Elo score, highest first. Scores are as published by the model’s provider or a public leaderboard. Where a hosted variant has no score of its own, it inherits the base model’s.
  • Limited to efficient models.
  • Includes models marked available or preview; retired and deprecated SKUs are excluded.
  • Prices are the rates each provider publishes, not estimates. How the blended rate is calculated.

58 otherwise-eligible models were left out because we hold no published value for the figure this list ranks by. They are excluded rather than assumed.

Frequently asked questions

What is a small language model?

A model with far fewer parameters than a frontier system, tuned for speed and cost rather than maximum capability. In practice it means lower latency, a much lower price per token, and sometimes the option to self-host, at the cost of accuracy on the hardest reasoning and coding problems.

What is LMArena Elo?

A rating derived from blind head-to-head comparisons: users are shown two anonymous model responses to the same prompt and pick the better one, and those votes feed an Elo rating like the one used in chess. It captures general helpfulness as people actually perceive it, rather than performance on a fixed test set.

Related rankings

Price it for your workload

Rankings are recomputed on every deploy from the catalogue. Prices checked 12 August 2026.