LMArena Leaderboard
Crowd-sourced head-to-head preference Elo rating. · Prices checked 12 August 2026
LMArena rates models from blind head-to-head comparisons: a user sends one prompt, sees two anonymous responses, and picks the better one. Those votes feed an Elo rating of the kind used in chess.
Because it aggregates human preference over open-ended prompts rather than performance on a fixed test set, it captures general helpfulness in a way scored benchmarks cannot, and it is correspondingly harder to game.
Leaderboard
60 of 150 scoredGPT-5.6 Sol leads on LMArena Elo at 1481 Elo. The median across the 150 models we hold a score for is 1248 Elo. The cheapest model on this board is Llama 4 Scout at $0.175/M blended, scoring 1265 Elo.
| # | Model | Provider | Score | Input /1M | Output /1M | Blended 70/30 | Context |
|---|---|---|---|---|---|---|---|
| 1 | GPT-5.6 Sol | OpenAI | 1481 Elo | $5.00 | $30.00 | $12.50 | 1,050,000 |
| 2 | GPT-5.6 Terra | OpenAI | 1465 Elo | $2.00 | $12.00 | $5.00 | 1,050,000 |
| 3 | GPT-5.6 Luna | OpenAI | 1451 Elo | $0.200 | $1.20 | $0.500 | 1,050,000 |
| 4 | GPT-5 | OpenAI | 1399 Elo | $1.25 | $10.00 | $3.88 | 400,000 |
| 5 | Claude Fable 5 | Anthropic | 1392 Elo | $10.00 | $50.00 | $22.00 | 1,000,000 |
| 6 | Gemini 2.5 Pro | 1389 Elo | $1.25 | $10.00 | $3.88 | 1,000,000 | |
| 7 | o3 | OpenAI | 1380 Elo | $10.00 | $40.00 | $19.00 | 200,000 |
| 8 | Claude Opus 4.6 | Anthropic | 1378 Elo | $5.00 | $25.00 | $11.00 | 1,000,000 |
| 9 | DeepSeek-R1-0528 | DeepSeek | 1370 Elo | $0.550 | $2.19 | $1.04 | 64,000 |
| 10 | Claude Opus 4.5 | Anthropic | 1366 Elo | $15.00 | $75.00 | $33.00 | 200,000 |
| 11 | DeepSeek-R1 | DeepSeek | 1361 Elo | $0.550 | $2.19 | $1.04 | 64,000 |
| 12 | DeepSeek-R1 (Groq) | Groq | 1361 Elo | $0.750 | $0.990 | $0.822 | 128,000 |
| 13 | DeepSeek-R1 (Together) | Together AI | 1361 Elo | $3.00 | $7.00 | $4.20 | 64,000 |
| 14 | DeepSeek-R1 (fw) | Fireworks AI | 1361 Elo | $3.00 | $8.00 | $4.50 | 64,000 |
| 15 | DeepSeek-R1 (DI) | DeepInfra | 1361 Elo | $0.550 | $2.19 | $1.04 | 64,000 |
| 16 | DeepSeek-R1 (Cerebras) | Cerebras | 1361 Elo | $0.550 | $0.990 | $0.682 | 64,000 |
| 17 | o1 | OpenAI | 1355 Elo | $15.00 | $60.00 | $28.50 | 200,000 |
| 18 | Claude Sonnet 4.6 | Anthropic | 1352 Elo | $3.00 | $15.00 | $6.60 | 1,000,000 |
| 19 | GPT-4.1 | OpenAI | 1346 Elo | $2.00 | $8.00 | $3.80 | 1,000,000 |
| 20 | Claude Sonnet 4.5 | Anthropic | 1340 Elo | $3.00 | $15.00 | $6.60 | 200,000 |
| 21 | o4-mini | OpenAI | 1340 Elo | $1.10 | $4.40 | $2.09 | 200,000 |
| 22 | Grok 3 | xAI | 1332 Elo | $3.00 | $15.00 | $6.60 | 131,000 |
| 23 | Grok 3 Fast | xAI | 1332 Elo | $5.00 | $25.00 | $11.00 | 131,000 |
| 24 | Gemini 2.5 Flash | 1325 Elo | $0.150 | $0.600 | $0.285 | 1,000,000 | |
| 25 | DeepSeek-V3-0324 | DeepSeek | 1325 Elo | $0.270 | $1.10 | $0.519 | 128,000 |
| 26 | Claude Sonnet 3.7 | Anthropic | 1320 Elo | $3.00 | $15.00 | $6.60 | 200,000 |
| 27 | DeepSeek-V3 | DeepSeek | 1318 Elo | $0.270 | $1.10 | $0.519 | 64,000 |
| 28 | GPT-4o | OpenAI | 1316 Elo | $2.50 | $10.00 | $4.75 | 128,000 |
| 29 | Llama 4 Maverick | Meta | 1310 Elo | $0.200 | $0.600 | $0.320 | 128,000 |
| 30 | o3-mini | OpenAI | 1305 Elo | $1.10 | $4.40 | $2.09 | 200,000 |
| 31 | Gemini 1.5 Pro | 1305 Elo | $1.25 | $5.00 | $2.38 | 2,000,000 | |
| 32 | Qwen2.5-Max | Qwen | 1296 Elo | $1.60 | $6.40 | $3.04 | 32,000 |
| 33 | GPT-4.1 mini | OpenAI | 1290 Elo | $0.400 | $1.60 | $0.760 | 1,000,000 |
| 34 | Grok 2 | xAI | 1290 Elo | $2.00 | $10.00 | $4.40 | 131,000 |
| 35 | Grok 2 Vision | xAI | 1290 Elo | $2.00 | $10.00 | $4.40 | 32,000 |
| 36 | Llama 3.1 405B | Meta | 1290 Elo | $0.800 | $0.800 | $0.800 | 128,000 |
| 37 | Llama 3.1 405B (Groq) | Groq | 1290 Elo | $2.99 | $2.99 | $2.99 | 128,000 |
| 38 | Llama 3.1 405B (Together) | Together AI | 1290 Elo | $3.50 | $3.50 | $3.50 | 128,000 |
| 39 | Llama 3.1 405B (fw) | Fireworks AI | 1290 Elo | $3.00 | $3.00 | $3.00 | 131,000 |
| 40 | Llama 3.1 405B (Rep) | Replicate | 1290 Elo | $9.50 | $9.50 | $9.50 | 128,000 |
| 41 | Llama 3.1 405B (DI) | DeepInfra | 1290 Elo | $0.800 | $0.800 | $0.800 | 128,000 |
| 42 | Llama-3.1-405B (Cerebras) | Cerebras | 1290 Elo | $0.990 | $0.990 | $0.990 | 128,000 |
| 43 | Llama 3.1 405B (Hyp) | Hyperbolic | 1290 Elo | $4.00 | $4.00 | $4.00 | 128,000 |
| 44 | Claude Sonnet 3.5 | Anthropic | 1283 Elo | $3.00 | $15.00 | $6.60 | 200,000 |
| 45 | Mistral Large 3 | Mistral | 1283 Elo | $0.500 | $1.50 | $0.800 | 256,000 |
| 46 | Mistral Large 2 | Mistral | 1283 Elo | $2.00 | $6.00 | $3.20 | 128,000 |
| 47 | Mistral-Large-2 (NIM) | Nvidia | 1283 Elo | $2.00 | $6.00 | $3.20 | 128,000 |
| 48 | Claude Haiku 4.5 | Anthropic | 1280 Elo | $0.800 | $4.00 | $1.76 | 200,000 |
| 49 | Nova Premier | Amazon | 1280 Elo | $2.50 | $12.50 | $5.50 | 300,000 |
| 50 | Command A | Cohere | 1280 Elo | $2.50 | $10.00 | $4.75 | 256,000 |
| 51 | GPT-4o mini | OpenAI | 1273 Elo | $0.150 | $0.600 | $0.285 | 128,000 |
| 52 | Gemini 2.0 Flash | 1270 Elo | $0.100 | $0.400 | $0.190 | 1,000,000 | |
| 53 | Llama-3.1-Nemotron-70B | Nvidia | 1267 Elo | $0.350 | $0.400 | $0.365 | 128,000 |
| 54 | Grok 3 mini | xAI | 1265 Elo | $0.300 | $0.500 | $0.360 | 131,000 |
| 55 | Llama 4 Scout | Meta | 1265 Elo | $0.100 | $0.350 | $0.175 | 512,000 |
| 56 | Pixtral Large | Mistral | 1262 Elo | $2.00 | $6.00 | $3.20 | 128,000 |
| 57 | QwQ-32B | Qwen | 1260 Elo | $0.600 | $2.40 | $1.14 | 131,000 |
| 58 | GPT-4 Turbo | OpenAI | 1257 Elo | $10.00 | $30.00 | $16.00 | 128,000 |
| 59 | Llama 3.3 70B | Meta | 1257 Elo | $0.230 | $0.400 | $0.281 | 128,000 |
| 60 | Llama 3.3 70B (Groq) | Groq | 1257 Elo | $0.590 | $0.790 | $0.650 | 128,000 |
About these scores
- Crowd-sourced head-to-head preference Elo rating.
- Scores are as published by the model’s provider or a public leaderboard. Provider-reported figures are self-reported and are not independently re-run by us.
- Where a hosted variant of an open-weights model has no score of its own, it inherits the base model’s — the weights are the same.
- Prices beside each score are the provider’s published token rates. How the blended rate is calculated.
Showing the top 60 of 150 models we hold a LMArena Elo score for.
Frequently asked questions
What is LMArena Elo?
A rating derived from blind pairwise votes. Users compare two anonymous model responses to the same prompt and choose the better one; the winner gains rating points and the loser gives them up, exactly as in chess. A higher Elo means users preferred that model more often across many head-to-head matchups.
Why does LMArena disagree with other benchmarks?
It measures a different thing. Scored benchmarks test correctness on problems with known answers; LMArena measures which response people prefer on open-ended prompts, which rewards tone, formatting, instruction-following and usefulness. A model can be excellent at graduate physics and still lose votes for being terse or badly formatted.
Other leaderboards
Rankings that weigh price too
Leaderboards are recomputed on every deploy from the catalogue. Prices checked 12 August 2026.