Best Small AI Models
Ranked from provider-published pricing · Prices checked 12 August 2026
Small models trade some peak capability for much lower cost and latency. This list ranks the efficient tier by LMArena Elo, a crowd-sourced rating from blind head-to-head comparisons, so the ordering reflects human preference rather than a single test.
The best of these land close to frontier models on everyday tasks while costing a fraction as much, which is what makes them the default choice for high-volume production work.
The ranking
top 20 of 69GPT-5.6 Luna from OpenAI leads this ranking at 1451 Elo. The median across the 69 qualifying models is 1184 Elo.
| # | Model | Provider | Efficient Models Ranked by LMArena Elo | MMLU | Input /1M | Output /1M | Blended 70/30 |
|---|---|---|---|---|---|---|---|
| 1 | GPT-5.6 Luna | OpenAI | 1451 Elo | — | $0.200 | $1.20 | $0.500 |
| 2 | GPT-4.1 mini | OpenAI | 1290 Elo | 81 | $0.400 | $1.60 | $0.760 |
| 3 | Claude Haiku 4.5 | Anthropic | 1280 Elo | 78 | $0.800 | $4.00 | $1.76 |
| 4 | GPT-4o mini | OpenAI | 1273 Elo | 82 | $0.150 | $0.600 | $0.285 |
| 5 | Gemini 2.0 Flash | 1270 Elo | 78 | $0.100 | $0.400 | $0.190 | |
| 6 | Grok 3 mini | xAI | 1265 Elo | 79 | $0.300 | $0.500 | $0.360 |
| 7 | Llama 4 Scout | Meta | 1265 Elo | 80 | $0.100 | $0.350 | $0.175 |
| 8 | Llama 3.3 70B | Meta | 1257 Elo | 86 | $0.230 | $0.400 | $0.281 |
| 9 | Llama 3.3 70B (Groq) | Groq | 1257 Elo | 86 | $0.590 | $0.790 | $0.650 |
| 10 | Llama 3.3 70B (Together) | Together AI | 1257 Elo | 86 | $0.880 | $0.880 | $0.880 |
| 11 | Llama 3.3 70B (DI) | DeepInfra | 1257 Elo | 86 | $0.230 | $0.400 | $0.281 |
| 12 | Llama-3.3-70B (Cerebras) | Cerebras | 1257 Elo | 86 | $0.590 | $0.990 | $0.710 |
| 13 | Llama 3.1 70B | Meta | 1248 Elo | 83 | $0.350 | $0.400 | $0.365 |
| 14 | Llama 3.1 70B (Groq) | Groq | 1248 Elo | 83 | $0.590 | $0.790 | $0.650 |
| 15 | Llama 3.1 70B (DI) | DeepInfra | 1248 Elo | 83 | $0.350 | $0.400 | $0.365 |
| 16 | Llama 3.1 70B (Hyp) | Hyperbolic | 1248 Elo | 83 | $0.400 | $0.400 | $0.400 |
| 17 | Llama 3.1 70B (Anyscale) | Anyscale | 1248 Elo | 83 | $1.00 | $1.00 | $1.00 |
| 18 | Llama 3.1 70B (Lepton) | Lepton AI | 1248 Elo | 83 | $0.800 | $0.800 | $0.800 |
| 19 | Llama 3.1 70B (DB) | Databricks | 1248 Elo | 83 | $1.00 | $3.00 | $1.60 |
| 20 | Yi-Lightning | 01.AI | 1244 Elo | 80 | $0.140 | $0.140 | $0.140 |
How this list is built
- Ranked by LMArena Elo score, highest first. Scores are as published by the model’s provider or a public leaderboard. Where a hosted variant has no score of its own, it inherits the base model’s.
- Limited to efficient models.
- Includes models marked available or preview; retired and deprecated SKUs are excluded.
- Prices are the rates each provider publishes, not estimates. How the blended rate is calculated.
58 otherwise-eligible models were left out because we hold no published value for the figure this list ranks by. They are excluded rather than assumed.
Frequently asked questions
What is a small language model?
A model with far fewer parameters than a frontier system, tuned for speed and cost rather than maximum capability. In practice it means lower latency, a much lower price per token, and sometimes the option to self-host, at the cost of accuracy on the hardest reasoning and coding problems.
What is LMArena Elo?
A rating derived from blind head-to-head comparisons: users are shown two anonymous model responses to the same prompt and pick the better one, and those votes feed an Elo rating like the one used in chess. It captures general helpfulness as people actually perceive it, rather than performance on a fixed test set.
Related rankings
Price it for your workload
Rankings are recomputed on every deploy from the catalogue. Prices checked 12 August 2026.