Claude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/MClaude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/M

LMArena Leaderboard

Crowd-sourced head-to-head preference Elo rating. · Prices checked 12 August 2026

LMArena rates models from blind head-to-head comparisons: a user sends one prompt, sees two anonymous responses, and picks the better one. Those votes feed an Elo rating of the kind used in chess.

Because it aggregates human preference over open-ended prompts rather than performance on a fixed test set, it captures general helpfulness in a way scored benchmarks cannot, and it is correspondingly harder to game.

Leaderboard

60 of 150 scored

GPT-5.6 Sol leads on LMArena Elo at 1481 Elo. The median across the 150 models we hold a score for is 1248 Elo. The cheapest model on this board is Llama 4 Scout at $0.175/M blended, scoring 1265 Elo.

#ModelProviderScoreInput /1MOutput /1MBlended 70/30Context
1GPT-5.6 SolOpenAI1481 Elo$5.00$30.00$12.501,050,000
2GPT-5.6 TerraOpenAI1465 Elo$2.00$12.00$5.001,050,000
3GPT-5.6 LunaOpenAI1451 Elo$0.200$1.20$0.5001,050,000
4GPT-5OpenAI1399 Elo$1.25$10.00$3.88400,000
5Claude Fable 5Anthropic1392 Elo$10.00$50.00$22.001,000,000
6Gemini 2.5 ProGoogle1389 Elo$1.25$10.00$3.881,000,000
7o3OpenAI1380 Elo$10.00$40.00$19.00200,000
8Claude Opus 4.6Anthropic1378 Elo$5.00$25.00$11.001,000,000
9DeepSeek-R1-0528DeepSeek1370 Elo$0.550$2.19$1.0464,000
10Claude Opus 4.5Anthropic1366 Elo$15.00$75.00$33.00200,000
11DeepSeek-R1DeepSeek1361 Elo$0.550$2.19$1.0464,000
12DeepSeek-R1 (Groq)Groq1361 Elo$0.750$0.990$0.822128,000
13DeepSeek-R1 (Together)Together AI1361 Elo$3.00$7.00$4.2064,000
14DeepSeek-R1 (fw)Fireworks AI1361 Elo$3.00$8.00$4.5064,000
15DeepSeek-R1 (DI)DeepInfra1361 Elo$0.550$2.19$1.0464,000
16DeepSeek-R1 (Cerebras)Cerebras1361 Elo$0.550$0.990$0.68264,000
17o1OpenAI1355 Elo$15.00$60.00$28.50200,000
18Claude Sonnet 4.6Anthropic1352 Elo$3.00$15.00$6.601,000,000
19GPT-4.1OpenAI1346 Elo$2.00$8.00$3.801,000,000
20Claude Sonnet 4.5Anthropic1340 Elo$3.00$15.00$6.60200,000
21o4-miniOpenAI1340 Elo$1.10$4.40$2.09200,000
22Grok 3xAI1332 Elo$3.00$15.00$6.60131,000
23Grok 3 FastxAI1332 Elo$5.00$25.00$11.00131,000
24Gemini 2.5 FlashGoogle1325 Elo$0.150$0.600$0.2851,000,000
25DeepSeek-V3-0324DeepSeek1325 Elo$0.270$1.10$0.519128,000
26Claude Sonnet 3.7Anthropic1320 Elo$3.00$15.00$6.60200,000
27DeepSeek-V3DeepSeek1318 Elo$0.270$1.10$0.51964,000
28GPT-4oOpenAI1316 Elo$2.50$10.00$4.75128,000
29Llama 4 MaverickMeta1310 Elo$0.200$0.600$0.320128,000
30o3-miniOpenAI1305 Elo$1.10$4.40$2.09200,000
31Gemini 1.5 ProGoogle1305 Elo$1.25$5.00$2.382,000,000
32Qwen2.5-MaxQwen1296 Elo$1.60$6.40$3.0432,000
33GPT-4.1 miniOpenAI1290 Elo$0.400$1.60$0.7601,000,000
34Grok 2xAI1290 Elo$2.00$10.00$4.40131,000
35Grok 2 VisionxAI1290 Elo$2.00$10.00$4.4032,000
36Llama 3.1 405BMeta1290 Elo$0.800$0.800$0.800128,000
37Llama 3.1 405B (Groq)Groq1290 Elo$2.99$2.99$2.99128,000
38Llama 3.1 405B (Together)Together AI1290 Elo$3.50$3.50$3.50128,000
39Llama 3.1 405B (fw)Fireworks AI1290 Elo$3.00$3.00$3.00131,000
40Llama 3.1 405B (Rep)Replicate1290 Elo$9.50$9.50$9.50128,000
41Llama 3.1 405B (DI)DeepInfra1290 Elo$0.800$0.800$0.800128,000
42Llama-3.1-405B (Cerebras)Cerebras1290 Elo$0.990$0.990$0.990128,000
43Llama 3.1 405B (Hyp)Hyperbolic1290 Elo$4.00$4.00$4.00128,000
44Claude Sonnet 3.5Anthropic1283 Elo$3.00$15.00$6.60200,000
45Mistral Large 3Mistral1283 Elo$0.500$1.50$0.800256,000
46Mistral Large 2Mistral1283 Elo$2.00$6.00$3.20128,000
47Mistral-Large-2 (NIM)Nvidia1283 Elo$2.00$6.00$3.20128,000
48Claude Haiku 4.5Anthropic1280 Elo$0.800$4.00$1.76200,000
49Nova PremierAmazon1280 Elo$2.50$12.50$5.50300,000
50Command ACohere1280 Elo$2.50$10.00$4.75256,000
51GPT-4o miniOpenAI1273 Elo$0.150$0.600$0.285128,000
52Gemini 2.0 FlashGoogle1270 Elo$0.100$0.400$0.1901,000,000
53Llama-3.1-Nemotron-70BNvidia1267 Elo$0.350$0.400$0.365128,000
54Grok 3 minixAI1265 Elo$0.300$0.500$0.360131,000
55Llama 4 ScoutMeta1265 Elo$0.100$0.350$0.175512,000
56Pixtral LargeMistral1262 Elo$2.00$6.00$3.20128,000
57QwQ-32BQwen1260 Elo$0.600$2.40$1.14131,000
58GPT-4 TurboOpenAI1257 Elo$10.00$30.00$16.00128,000
59Llama 3.3 70BMeta1257 Elo$0.230$0.400$0.281128,000
60Llama 3.3 70B (Groq)Groq1257 Elo$0.590$0.790$0.650128,000

About these scores

  • Crowd-sourced head-to-head preference Elo rating.
  • Scores are as published by the model’s provider or a public leaderboard. Provider-reported figures are self-reported and are not independently re-run by us.
  • Where a hosted variant of an open-weights model has no score of its own, it inherits the base model’s — the weights are the same.
  • Prices beside each score are the provider’s published token rates. How the blended rate is calculated.

Showing the top 60 of 150 models we hold a LMArena Elo score for.

Frequently asked questions

What is LMArena Elo?

A rating derived from blind pairwise votes. Users compare two anonymous model responses to the same prompt and choose the better one; the winner gains rating points and the loser gives them up, exactly as in chess. A higher Elo means users preferred that model more often across many head-to-head matchups.

Why does LMArena disagree with other benchmarks?

It measures a different thing. Scored benchmarks test correctness on problems with known answers; LMArena measures which response people prefer on open-ended prompts, which rewards tone, formatting, instruction-following and usefulness. A model can be excellent at graduate physics and still lose votes for being terse or badly formatted.

Other leaderboards

Rankings that weigh price too

Leaderboards are recomputed on every deploy from the catalogue. Prices checked 12 August 2026.