Index · Industry Metrics
AI model pricing is charged per million tokens, split into input (the prompt) and output (the completion) rates. Tokenando tracks live pricing for 678 models across 68 providers, with a blended 70/30 reference rate for quick comparison. Compare input, output, and blended costs in the table below. How this is calculated →
As of August 2026, half of the models on this index blend to between $0.27 and $2.30 per million tokens, with a median of $0.80. The cheapest paid model is $0.020 and the most expensive is OpenAI o1-pro at $285.00, while 10 models carry no per-token charge at all. Blended figures are our own 70/30 weighting of the published input and output rates, not a price any provider quotes.
| Provider↕ | Model↕ | Input↕ | Output↕ | Context↕ | Blend · 70/30↑ |
|---|---|---|---|---|---|
| Lyria 3 Pro Preview google/lyria-3-pro-preview | $0/M | $0/M | 1.0M | $0/M | |
| Lyria 3 Clip Preview google/lyria-3-clip-preview | $0/M | $0/M | 1.0M | $0/M | |
| Ox Alpha stealth/ox-alpha | $0/M | $0/M | 1.0M | $0/M | |
| Dots3-Note Preview (free) dots-studio/dots-3-note-preview:free | $0/M | $0/M | 512K | $0/M | |
| North Mini Code (free) cohere/north-mini-code:free | $0/M | $0/M | 256K | $0/M | |
| Nemotron 3 Nano Omni (free) nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free | $0/M | $0/M | 256K | $0/M | |
| CoBuddy (free) baidu/cobuddy:free | $0/M | $0/M | 131K | $0/M | |
| Nemotron 3.5 Content Safety (free) nvidia/nemotron-3.5-content-safety:free | $0/M | $0/M | 128K | $0/M | |
| LiquidAI: LFM2.5-2.6B (free) liquid/lfm-2.5-2.6b:free | $0/M | $0/M | 66K | $0/M | |
| Gemma 3n 2B (free) google/gemma-3n-e2b-it:free | $0/M | $0/M | 8K | $0/M | |
| Whisper Large v3 Whisper Large v3 | $0.0036/M | — | — | $0.0036/M | |
| Gemma 2 2B Gemma 2 2B | $0.0200/M | $0.0200/M | 8K | $0.0200/M | |
| Titan Embeddings V2 Titan Embeddings V2 | $0.0200/M | — | 8K | $0.0200/M | |
| text-embedding-3-small text-embedding-3-small | $0.0200/M | — | 8K | $0.0200/M | |
| Mistral Nemo mistralai/mistral-nemo | $0.0190/M | $0.0300/M | 131K | $0.0223/M | |
| text-embedding-004 text-embedding-004 | $0.0250/M | — | 2K | $0.0250/M | |
| Llama 3.1 8B Instruct meta-llama/llama-3.1-8b-instruct | $0.0200/M | $0.0500/M | 16K | $0.0290/M | |
| Qwen2.5-1.5B-Instruct Qwen2.5-1.5B-Instruct | $0.0300/M | $0.0300/M | 32K | $0.0300/M | |
| Llama 3 8B Instruct meta-llama/llama-3-8b-instruct | $0.0300/M | $0.0400/M | 8K | $0.0330/M | |
| Ling-3.0-flash inclusionai/ling-3.0-flash | $0.0210/M | $0.0630/M | 262K | $0.0336/M | |
| Llama 3.2 1B Llama 3.2 1B | $0.0400/M | $0.0400/M | 128K | $0.0400/M | |
| Phi-3-Mini-4K (NIM) Phi-3-Mini-4K (NIM) | $0.0400/M | $0.0400/M | 4K | $0.0400/M | |
| SDXL (Rep) SDXL (Rep) | $0.0400/M | — | — | $0.0400/M | |
| Llama 3 8B Lunaris sao10k/l3-lunaris-8b | $0.0400/M | $0.0500/M | 8K | $0.0430/M | |
| Granite 4.0 Micro ibm-granite/granite-4.0-h-micro | $0.0170/M | $0.1120/M | 131K | $0.0455/M | |
| Nex-N2-Mini nex-agi/nex-n2-mini | $0.0250/M | $0.1000/M | 262K | $0.0475/M | |
| Qwen2.5-3B-Instruct Qwen2.5-3B-Instruct | $0.0500/M | $0.0500/M | 32K | $0.0500/M | |
| ABAB 5.5c ABAB 5.5c | $0.0500/M | $0.0500/M | 16K | $0.0500/M | |
| Granite 3.1 2B Instruct Granite 3.1 2B Instruct | $0.0300/M | $0.1000/M | 128K | $0.0510/M | |
| Doubao-Lite-128k Doubao-Lite-128k | $0.0400/M | $0.0900/M | 128K | $0.0550/M | |
| Doubao-Lite-32k Doubao-Lite-32k | $0.0400/M | $0.0900/M | 32K | $0.0550/M | |
| Solar Pro 4 upstage/solar-pro4 | $0.0300/M | $0.1200/M | 524K | $0.0570/M | |
| Qwen2.5 7B Instruct qwen/qwen-2.5-7b-instruct | $0.0400/M | $0.1000/M | 33K | $0.0580/M | |
| Llama 3.1 8B (Groq) Llama 3.1 8B (Groq) | $0.0500/M | $0.0800/M | 128K | $0.0590/M | |
| Mistral Small 3 mistralai/mistral-small-24b-instruct-2501 | $0.0500/M | $0.0800/M | 33K | $0.0590/M | |
| Qwen3.7 Flash qwen/qwen3.7-flash | $0.0300/M | $0.1300/M | 1M | $0.0600/M | |
| Llama 3.2 3B Llama 3.2 3B | $0.0600/M | $0.0600/M | 128K | $0.0600/M | |
| Llama 3.2 3B (Groq) Llama 3.2 3B (Groq) | $0.0600/M | $0.0600/M | 128K | $0.0600/M | |
| MythoMax 13B gryphe/mythomax-l2-13b | $0.0600/M | $0.0600/M | 8K | $0.0600/M | |
| Gemma 2 9B (DI) Gemma 2 9B (DI) | $0.0600/M | $0.0600/M | 8K | $0.0600/M | |
| Qwen-Turbo qwen/qwen-turbo | $0.0325/M | $0.1300/M | 131K | $0.0617/M | |
| gpt-oss-20b openai/gpt-oss-20b | $0.0300/M | $0.1400/M | 131K | $0.0630/M | |
| Gemma 3 4B google/gemma-3-4b-it | $0.0500/M | $0.1000/M | 131K | $0.0650/M | |
| Granite 4.1 8B ibm-granite/granite-4.1-8b | $0.0500/M | $0.1000/M | 131K | $0.0650/M | |
| voyage-3-lite voyage-3-lite | $0.0650/M | — | 32K | $0.0650/M | |
| Nova Micro Nova Micro | $0.0350/M | $0.1400/M | 128K | $0.0665/M | |
| Nova Micro 1.0 amazon/nova-micro-v1 | $0.0350/M | $0.1400/M | 128K | $0.0665/M | |
| Mistral 7B (DI) Mistral 7B (DI) | $0.0700/M | $0.0700/M | 32K | $0.0700/M | |
| Mistral 7B (Lepton) Mistral 7B (Lepton) | $0.0700/M | $0.0700/M | 32K | $0.0700/M | |
| Gemini 1.5 Flash-8B Gemini 1.5 Flash-8B | $0.0375/M | $0.1500/M | 1M | $0.0712/M | |
| Command R7B (12-2024) cohere/command-r7b-12-2024 | $0.0375/M | $0.1500/M | 128K | $0.0712/M | |
| Nemotron Nano 9B V2 nvidia/nemotron-nano-9b-v2 | $0.0400/M | $0.1600/M | 131K | $0.0760/M | |
| Phi-4 Mini Phi-4 Mini | $0.0400/M | $0.1600/M | 128K | $0.0760/M | |
| gpt-oss-120b openai/gpt-oss-120b | $0.0370/M | $0.1700/M | 131K | $0.0769/M | |
| GPT-5 Nano (batch) openai/gpt-5-nano:batch | $0.0250/M | $0.2000/M | 400K | $0.0775/M | |
| Laguna XS 2.1 poolside/laguna-xs-2.1 | $0.0600/M | $0.1200/M | 262K | $0.0780/M | |
| Gemma 3n 4B google/gemma-3n-e4b-it | $0.0600/M | $0.1200/M | 33K | $0.0780/M | |
| Llama 3.2 1B Instruct meta-llama/llama-3.2-1b-instruct | $0.0270/M | $0.2010/M | 60K | $0.0792/M | |
| Gemma 3 12B google/gemma-3-12b-it | $0.0500/M | $0.1500/M | 131K | $0.0800/M | |
| Mistral 7B v0.3 Mistral 7B v0.3 | $0.0800/M | $0.0800/M | 32K | $0.0800/M | |
| Nova Canvas Nova Canvas | $0.0800/M | — | — | $0.0800/M | |
| Hy-MT2-1.8B tencent/hy-mt2-1.8b | $0.0440/M | $0.1770/M | 8K | $0.0839/M | |
| Gemma 3 9B Gemma 3 9B | $0.0900/M | $0.0900/M | 128K | $0.0900/M | |
| Gemma 2 9B Gemma 2 9B | $0.0900/M | $0.0900/M | 8K | $0.0900/M | |
| Phi 4 microsoft/phi-4 | $0.0700/M | $0.1400/M | 16K | $0.0910/M | |
| Qwen3 235B A22B Instruct 2507 qwen/qwen3-235b-a22b-2507 | $0.0900/M | $0.1000/M | 262K | $0.0930/M | |
| Gemini 2.5 Flash Lite (batch) google/gemini-2.5-flash-lite:batch | $0.0500/M | $0.2000/M | 1.0M | $0.0950/M | |
| GPT-4.1 Nano (batch) openai/gpt-4.1-nano:batch | $0.0500/M | $0.2000/M | 1.0M | $0.0950/M | |
| Nemotron 3 Nano 30B A3B nvidia/nemotron-3-nano-30b-a3b | $0.0500/M | $0.2000/M | 262K | $0.0950/M | |
| Ministral 3 3B 2512 mistralai/ministral-3b-2512 | $0.1000/M | $0.1000/M | 131K | $0.1000/M | |
| Llama 3.1 8B (fw) Llama 3.1 8B (fw) | $0.1000/M | $0.1000/M | 131K | $0.1000/M | |
| Llama 3.1 8B Llama 3.1 8B | $0.1000/M | $0.1000/M | 128K | $0.1000/M | |
| Llama-3.1-8B (Cerebras) Llama-3.1-8B (Cerebras) | $0.1000/M | $0.1000/M | 128K | $0.1000/M | |
| Qwen2.5-Coder-7B Qwen2.5-Coder-7B | $0.1000/M | $0.1000/M | 128K | $0.1000/M | |
| Llama 3.1 8B (Together) Llama 3.1 8B (Together) | $0.1000/M | $0.1000/M | 128K | $0.1000/M | |
| Reka Edge rekaai/reka-edge | $0.1000/M | $0.1000/M | 16K | $0.1000/M | |
| StableCode 3B StableCode 3B | $0.1000/M | $0.1000/M | 16K | $0.1000/M | |
| mistral-embed mistral-embed | $0.1000/M | — | 8K | $0.1000/M | |
| Solar Embedding Solar Embedding | $0.1000/M | — | 4K | $0.1000/M | |
| StableLM 2 1.6B StableLM 2 1.6B | $0.1000/M | $0.1000/M | 4K | $0.1000/M | |
| Embed v3 English Embed v3 English | $0.1000/M | — | 512 | $0.1000/M | |
| Embed v3 Multilingual Embed v3 Multilingual | $0.1000/M | — | 512 | $0.1000/M | |
| Gemma 3 27B google/gemma-3-27b-it | $0.0800/M | $0.1600/M | 131K | $0.1040/M | |
| Hy3 preview tencent/hy3-preview | $0.0630/M | $0.2100/M | 262K | $0.1071/M | |
| DeepSeek-R1-Distill-8B DeepSeek-R1-Distill-8B | $0.0700/M | $0.2000/M | 128K | $0.1090/M | |
| Granite 3.1 8B Instruct Granite 3.1 8B Instruct | $0.0500/M | $0.2500/M | 128K | $0.1100/M | |
| Llama 3 8B (Rep) Llama 3 8B (Rep) | $0.0500/M | $0.2500/M | 8K | $0.1100/M | |
| Granite 3.0 8B Dense Granite 3.0 8B Dense | $0.0500/M | $0.2500/M | 4K | $0.1100/M | |
| Ministral 8B mistralai/ministral-8b | $0.1100/M | $0.1100/M | 128K | $0.1100/M | |
| Mistral Small 3.2 24B mistralai/mistral-small-3.2-24b-instruct | $0.0750/M | $0.2000/M | 128K | $0.1125/M | |
| Nova Lite Nova Lite | $0.0600/M | $0.2400/M | 300K | $0.1140/M | |
| Nova Lite 1.0 amazon/nova-lite-v1 | $0.0600/M | $0.2400/M | 300K | $0.1140/M | |
| Qwen3 14B qwen/qwen3-14b | $0.0600/M | $0.2400/M | 41K | $0.1140/M | |
| Qwen3.5-9B qwen/qwen3.5-9b | $0.1000/M | $0.1500/M | 262K | $0.1150/M | |
| DeepSeek V4 Flash 0423 deepseek/deepseek-v4-flash | $0.0886/M | $0.1772/M | 1.0M | $0.1152/M | |
| Nemotron 3.5 Lightning nvidia/nemotron-3.5-lightning | $0.0800/M | $0.2000/M | 262K | $0.1160/M | |
| Laguna S 2.1 poolside/laguna-s-2.1 | $0.0900/M | $0.1800/M | 1.0M | $0.1170/M | |
| voyage-3 voyage-3 | $0.1200/M | — | 32K | $0.1200/M | |
| voyage-finance-2 voyage-finance-2 | $0.1200/M | — | 32K | $0.1200/M | |
| voyage-law-2 voyage-law-2 | $0.1200/M | — | 32K | $0.1200/M | |
| voyage-multilingual-2 voyage-multilingual-2 | $0.1200/M | — | 32K | $0.1200/M | |
| Moonshot v1 8k Moonshot v1 8k | $0.1200/M | $0.1200/M | 8K | $0.1200/M | |
| Qwen3.5-Flash qwen/qwen3.5-flash-02-23 | $0.0650/M | $0.2600/M | 1M | $0.1235/M | |
| Qwen3 32B qwen/qwen3-32b | $0.0800/M | $0.2400/M | 41K | $0.1280/M | |
| Muse Spark 1.2 Contributor meta/muse-spark-1.2-contributor | $0.1000/M | $0.2000/M | 1.0M | $0.1300/M | |
| UI-TARS 7B bytedance/ui-tars-1.5-7b | $0.1000/M | $0.2000/M | 128K | $0.1300/M | |
| Reka Flash 3 rekaai/reka-flash-3 | $0.1000/M | $0.2000/M | 66K | $0.1300/M | |
| text-embedding-3-large text-embedding-3-large | $0.1300/M | — | 8K | $0.1300/M | |
| Qwen3 Coder 30B A3B Instruct qwen/qwen3-coder-30b-a3b-instruct | $0.0700/M | $0.2800/M | 160K | $0.1330/M | |
| ERNIE 4.5 21B A3B Thinking baidu/ernie-4.5-21b-a3b-thinking | $0.0700/M | $0.2800/M | 131K | $0.1330/M | |
| ERNIE 4.5 21B A3B baidu/ernie-4.5-21b-a3b | $0.0700/M | $0.2800/M | 120K | $0.1330/M | |
| Llama 3.2 3B Instruct meta-llama/llama-3.2-3b-instruct | $0.0500/M | $0.3300/M | 80K | $0.1340/M | |
| Mistral 7B Instruct v0.1 mistralai/mistral-7b-instruct-v0.1 | $0.1100/M | $0.1900/M | 3K | $0.1340/M | |
| GLM-4-Air GLM-4-Air | $0.1400/M | $0.1400/M | 128K | $0.1400/M | |
| Qwen3 30B A3B qwen/qwen3-30b-a3b | $0.0800/M | $0.2800/M | 41K | $0.1400/M | |
| Yi-Lightning Yi-Lightning | $0.1400/M | $0.1400/M | 16K | $0.1400/M | |
| Hy-MT2-30B-A3B tencent/hy-mt2-30b-a3b | $0.0740/M | $0.2950/M | 8K | $0.1403/M | |
| Hy-MT2-7B tencent/hy-mt2-7b | $0.0740/M | $0.2950/M | 8K | $0.1403/M | |
| Gemini 2.0 Flash Lite google/gemini-2.0-flash-lite-001 | $0.0750/M | $0.3000/M | 1.0M | $0.1425/M | |
| Gemini 1.5 Flash Gemini 1.5 Flash | $0.0750/M | $0.3000/M | 1M | $0.1425/M | |
| Seed 1.6 Flash bytedance-seed/seed-1.6-flash | $0.0750/M | $0.3000/M | 262K | $0.1425/M | |
| gpt-oss-safeguard-20b openai/gpt-oss-safeguard-20b | $0.0750/M | $0.3000/M | 131K | $0.1425/M | |
| GPT-4o-mini (batch) openai/gpt-4o-mini:batch | $0.0750/M | $0.3000/M | 128K | $0.1425/M | |
| Ministral 3 8B 2512 mistralai/ministral-8b-2512 | $0.1500/M | $0.1500/M | 262K | $0.1500/M | |
| Pixtral 12B Pixtral 12B | $0.1500/M | $0.1500/M | 128K | $0.1500/M | |
| Mistral-NeMo-12B (NIM) Mistral-NeMo-12B (NIM) | $0.1500/M | $0.1500/M | 128K | $0.1500/M | |
| Mistral 7B (Anyscale) Mistral 7B (Anyscale) | $0.1500/M | $0.1500/M | 32K | $0.1500/M | |
| Yi-6B Yi-6B | $0.1500/M | $0.1500/M | 4K | $0.1500/M | |
| Gemma 4 26B A4B google/gemma-4-26b-a4b-it | $0.0700/M | $0.3400/M | 262K | $0.1510/M | |
| GPT-5 Nano openai/gpt-5-nano | $0.0500/M | $0.4000/M | 400K | $0.1550/M | |
| Qwen3 8B qwen/qwen3-8b | $0.0500/M | $0.4000/M | 41K | $0.1550/M | |
| Llama 4 Scout meta-llama/llama-4-scout | $0.1000/M | $0.3000/M | 328K | $0.1600/M | |
| Qwen3 30B A3B Instruct 2507 qwen/qwen3-30b-a3b-instruct-2507 | $0.1000/M | $0.3000/M | 262K | $0.1600/M | |
| Step 3.5 Flash stepfun/step-3.5-flash | $0.1000/M | $0.3000/M | 262K | $0.1600/M | |
| Devstral Small 1.1 mistralai/devstral-small | $0.1000/M | $0.3000/M | 131K | $0.1600/M | |
| Llama 3.2 11B Vision Llama 3.2 11B Vision | $0.1600/M | $0.1600/M | 128K | $0.1600/M | |
| Mistral Small 3.2 Mistral Small 3.2 | $0.1000/M | $0.3000/M | 128K | $0.1600/M | |
| Voxtral Small 24B 2507 mistralai/voxtral-small-24b-2507 | $0.1000/M | $0.3000/M | 32K | $0.1600/M | |
| Phi 4 Mini Instruct microsoft/phi-4-mini-instruct | $0.0800/M | $0.3500/M | 128K | $0.1610/M | |
| Doubao-Pro-32k Doubao-Pro-32k | $0.1100/M | $0.2800/M | 32K | $0.1610/M | |
| Z.ai: GLM 4.7 Flash z-ai/glm-4.7-flash | $0.0600/M | $0.4000/M | 203K | $0.1620/M | |
| Titan Text Lite Titan Text Lite | $0.1500/M | $0.2000/M | 4K | $0.1650/M | |
| Llama 3.3 70B Instruct meta-llama/llama-3.3-70b-instruct | $0.1000/M | $0.3200/M | 131K | $0.1660/M | |
| DeepSeek-R1-Distill-14B DeepSeek-R1-Distill-14B | $0.1000/M | $0.3500/M | 128K | $0.1750/M | |
| Qwen3 30B A3B Thinking 2507 qwen/qwen3-30b-a3b-thinking-2507 | $0.0800/M | $0.4000/M | 131K | $0.1760/M | |
| Nemotron 3 Super nvidia/nemotron-3-super-120b-a12b | $0.0850/M | $0.4000/M | 262K | $0.1795/M | |
| Llama Guard 4 12B meta-llama/llama-guard-4-12b | $0.1800/M | $0.1800/M | 164K | $0.1800/M | |
| Llama 3.2 11B Vision (Groq) Llama 3.2 11B Vision (Groq) | $0.1800/M | $0.1800/M | 128K | $0.1800/M | |
| voyage-3-large voyage-3-large | $0.1800/M | — | 32K | $0.1800/M | |
| voyage-code-3 voyage-code-3 | $0.1800/M | — | 32K | $0.1800/M | |
| CodeLlama 7B Instruct CodeLlama 7B Instruct | $0.1800/M | $0.1800/M | 16K | $0.1800/M | |
| Llama 2 7B Chat Llama 2 7B Chat | $0.1800/M | $0.1800/M | 4K | $0.1800/M | |
| Falcon 7B Instruct Falcon 7B Instruct | $0.1800/M | $0.1800/M | 2K | $0.1800/M | |
| MiMo-V2.5 xiaomi/mimo-v2.5 | $0.1400/M | $0.2800/M | 1.1M | $0.1820/M | |
| DeepSeek V4 Flash deepseek/deepseek-v4-flash | $0.1400/M | $0.2800/M | 1.0M | $0.1820/M | |
| DeepSeek V4 Flash 0731 deepseek/deepseek-v4-flash-0731 | $0.1400/M | $0.2800/M | 1.0M | $0.1820/M | |
| DeepSeek-Coder-V2 DeepSeek-Coder-V2 | $0.1400/M | $0.2800/M | 128K | $0.1820/M | |
| Gemini 2.5 Flash Lite google/gemini-2.5-flash-lite | $0.1000/M | $0.4000/M | 1.0M | $0.1900/M | |
| Gemini 2.5 Flash Lite Preview 09-2025 google/gemini-2.5-flash-lite-preview-09-2025 | $0.1000/M | $0.4000/M | 1.0M | $0.1900/M | |
| GPT-4.1 Nano openai/gpt-4.1-nano | $0.1000/M | $0.4000/M | 1.0M | $0.1900/M | |
| Gemini 2.0 Flash google/gemini-2.0-flash-001 | $0.1000/M | $0.4000/M | 1M | $0.1900/M | |
| Seed-2.0-Mini bytedance-seed/seed-2.0-mini | $0.1000/M | $0.4000/M | 262K | $0.1900/M | |
| Llama 3.3 Nemotron Super 49B V1.5 nvidia/llama-3.3-nemotron-super-49b-v1.5 | $0.1000/M | $0.4000/M | 131K | $0.1900/M | |
| Qwen3 VL 32B Instruct qwen/qwen3-vl-32b-instruct | $0.1040/M | $0.4160/M | 131K | $0.1976/M | |
| Ministral 3 14B 2512 mistralai/ministral-14b-2512 | $0.2000/M | $0.2000/M | 262K | $0.2000/M | |
| Mixtral 8x7B (fw) Mixtral 8x7B (fw) | $0.2000/M | $0.2000/M | 32K | $0.2000/M | |
| Llama 3 8B Llama 3 8B | $0.2000/M | $0.2000/M | 8K | $0.2000/M | |
| Gemma 2 9B (Groq) Gemma 2 9B (Groq) | $0.2000/M | $0.2000/M | 8K | $0.2000/M | |
| Falcon 11B Falcon 11B | $0.2000/M | $0.2000/M | 8K | $0.2000/M | |
| Qwen3 VL 8B Instruct qwen/qwen3-vl-8b-instruct | $0.0800/M | $0.5000/M | 131K | $0.2060/M | |
| Nous: Hermes 4 70B nousresearch/hermes-4-70b | $0.1300/M | $0.4000/M | 131K | $0.2110/M | |
| Gemma 4 31B google/gemma-4-31b-it | $0.1400/M | $0.4000/M | 262K | $0.2180/M | |
| Qwen VL Plus qwen/qwen-vl-plus | $0.1365/M | $0.4095/M | 131K | $0.2184/M | |
| CodeLlama 13B Instruct CodeLlama 13B Instruct | $0.2200/M | $0.2200/M | 16K | $0.2200/M | |
| Llama 2 13B Chat Llama 2 13B Chat | $0.2200/M | $0.2200/M | 4K | $0.2200/M | |
| Mixtral 8x7B v0.1 Mixtral 8x7B v0.1 | $0.2400/M | $0.2400/M | 32K | $0.2400/M | |
| Mixtral 8x7B (Groq) Mixtral 8x7B (Groq) | $0.2400/M | $0.2400/M | 32K | $0.2400/M | |
| Mixtral 8x7B (DI) Mixtral 8x7B (DI) | $0.2400/M | $0.2400/M | 32K | $0.2400/M | |
| Llama 3.2 11B Vision Instruct meta-llama/llama-3.2-11b-vision-instruct | $0.2450/M | $0.2450/M | 131K | $0.2450/M | |
| DeepSeek V3.2 deepseek/deepseek-v3.2 | $0.2145/M | $0.3217/M | 131K | $0.2467/M | |
| Qwen3 VL 30B A3B Instruct qwen/qwen3-vl-30b-a3b-instruct | $0.1300/M | $0.5200/M | 131K | $0.2470/M | |
| Phi-3.5 Mini Phi-3.5 Mini | $0.1300/M | $0.5200/M | 128K | $0.2470/M | |
| Phi-3 Mini Phi-3 Mini | $0.1300/M | $0.5200/M | 128K | $0.2470/M | |
| GPT-5.6 Luna Pro (batch) openai/gpt-5.6-luna-pro:batch | $0.1000/M | $0.6000/M | 1.1M | $0.2500/M | |
| GPT-5.6 Luna (batch) openai/gpt-5.6-luna:batch | $0.1000/M | $0.6000/M | 1.1M | $0.2500/M | |
| Codestral Mamba Codestral Mamba | $0.2500/M | $0.2500/M | 256K | $0.2500/M | |
| Hy3 tencent/hy3 | $0.1320/M | $0.5280/M | 262K | $0.2508/M | |
| Olmo 3 32B Think allenai/olmo-3-32b-think | $0.1500/M | $0.5000/M | 66K | $0.2550/M | |
| GPT-5.4 Nano (batch) openai/gpt-5.4-nano:batch | $0.1000/M | $0.6250/M | 400K | $0.2575/M | |
| Jamba 1.5 Mini Jamba 1.5 Mini | $0.2000/M | $0.4000/M | 256K | $0.2600/M | |
| ERNIE 4.5 VL 28B A3B baidu/ernie-4.5-vl-28b-a3b | $0.1400/M | $0.5600/M | 30K | $0.2660/M | |
| Hunyuan A13B Instruct tencent/hunyuan-a13b-instruct | $0.1400/M | $0.5700/M | 131K | $0.2690/M | |
| Llama 3.3 70B Llama 3.3 70B | $0.2300/M | $0.4000/M | 128K | $0.2810/M | |
| Llama 3.3 70B (DI) Llama 3.3 70B (DI) | $0.2300/M | $0.4000/M | 128K | $0.2810/M | |
| Llama 4 Maverick meta-llama/llama-4-maverick | $0.1500/M | $0.6000/M | 1.0M | $0.2850/M | |
| Mistral Small 4 mistralai/mistral-small-2603 | $0.1500/M | $0.6000/M | 262K | $0.2850/M | |
| KAT-Coder-Air V2.5 kwaipilot/kat-coder-air-v2.5 | $0.1500/M | $0.6000/M | 256K | $0.2850/M | |
| GPT-4o-mini openai/gpt-4o-mini | $0.1500/M | $0.6000/M | 128K | $0.2850/M | |
| Phi-3 Small Phi-3 Small | $0.1500/M | $0.6000/M | 128K | $0.2850/M | |
| Solar Pro 3 upstage/solar-pro-3 | $0.1500/M | $0.6000/M | 128K | $0.2850/M | |
| Command R (08-2024) cohere/command-r-08-2024 | $0.1500/M | $0.6000/M | 128K | $0.2850/M | |
| GPT-4o-mini Search Preview openai/gpt-4o-mini-search-preview | $0.1500/M | $0.6000/M | 128K | $0.2850/M | |
| Solar Mini Solar Mini | $0.1500/M | $0.6000/M | 32K | $0.2850/M | |
| Grok 4.1 Fast x-ai/grok-4.1-fast | $0.2000/M | $0.5000/M | 2M | $0.2900/M | |
| Grok 4 Fast x-ai/grok-4-fast | $0.2000/M | $0.5000/M | 2M | $0.2900/M | |
| DeepSeek-R1-Distill-32B DeepSeek-R1-Distill-32B | $0.2000/M | $0.5000/M | 128K | $0.2900/M | |
| R1 Distill Qwen 32B deepseek/deepseek-r1-distill-qwen-32b | $0.2900/M | $0.2900/M | 33K | $0.2900/M | |
| Qwen2-VL-7B Qwen2-VL-7B | $0.3000/M | $0.3000/M | 32K | $0.3000/M | |
| Yi-9B Yi-9B | $0.3000/M | $0.3000/M | 4K | $0.3000/M | |
| Yi-VL-6B Yi-VL-6B | $0.3000/M | $0.3000/M | 4K | $0.3000/M | |
| StableLM 2 12B StableLM 2 12B | $0.3000/M | $0.3000/M | 4K | $0.3000/M | |
| Qwen3 Next 80B A3B Instruct qwen/qwen3-next-80b-a3b-instruct | $0.0975/M | $0.7800/M | 262K | $0.3022/M | |
| Qwen3 Next 80B A3B Thinking qwen/qwen3-next-80b-a3b-thinking | $0.0975/M | $0.7800/M | 131K | $0.3022/M | |
| Phi-3.5 MoE Phi-3.5 MoE | $0.1600/M | $0.6400/M | 128K | $0.3040/M | |
| DeepSeek V3.2 Exp deepseek/deepseek-v3.2-exp | $0.2700/M | $0.4100/M | 164K | $0.3120/M | |
| Gemini 3.1 Flash Lite (batch) google/gemini-3.1-flash-lite:batch | $0.1250/M | $0.7500/M | 1.0M | $0.3125/M | |
| Codestral Codestral | $0.2000/M | $0.6000/M | 256K | $0.3200/M | |
| Nemotron Nano 12B 2 VL nvidia/nemotron-nano-12b-v2-vl | $0.2000/M | $0.6000/M | 131K | $0.3200/M | |
| Saba mistralai/mistral-saba | $0.2000/M | $0.6000/M | 33K | $0.3200/M | |
| Phi-3 Medium Phi-3 Medium | $0.1700/M | $0.6800/M | 128K | $0.3230/M | |
| Qwen3 Coder Next qwen/qwen3-coder-next | $0.1200/M | $0.8000/M | 262K | $0.3240/M | |
| Rocinante 12B thedrummer/rocinante-12b | $0.2500/M | $0.5000/M | 66K | $0.3250/M | |
| DeepSeek V3.1 deepseek/deepseek-chat-v3.1 | $0.1500/M | $0.7500/M | 33K | $0.3300/M | |
| ABAB 5.5s ABAB 5.5s | $0.1500/M | $0.7500/M | 16K | $0.3300/M | |
| Z.ai: GLM 4.5 Air z-ai/glm-4.5-air | $0.1300/M | $0.8500/M | 131K | $0.3460/M | |
| Qwen2.5-14B-Instruct Qwen2.5-14B-Instruct | $0.3500/M | $0.3500/M | 128K | $0.3500/M | |
| Qwen2.5-Coder-14B Qwen2.5-Coder-14B | $0.3500/M | $0.3500/M | 128K | $0.3500/M | |
| CodeLlama 34B Instruct CodeLlama 34B Instruct | $0.3500/M | $0.3500/M | 16K | $0.3500/M | |
| Grok 3 Mini x-ai/grok-3-mini | $0.3000/M | $0.5000/M | 131K | $0.3600/M | |
| Grok 3 Mini Beta x-ai/grok-3-mini-beta | $0.3000/M | $0.5000/M | 131K | $0.3600/M | |
| Cydonia 24B V4.1 thedrummer/cydonia-24b-v4.1 | $0.3000/M | $0.5000/M | 131K | $0.3600/M | |
| Llama 3.1 70B Llama 3.1 70B | $0.3500/M | $0.4000/M | 128K | $0.3650/M | |
| Llama 3.1 70B (DI) Llama 3.1 70B (DI) | $0.3500/M | $0.4000/M | 128K | $0.3650/M | |
| Qwen2.5-72B (DI) Qwen2.5-72B (DI) | $0.3500/M | $0.4000/M | 128K | $0.3650/M | |
| Llama-3.1-Nemotron-70B Llama-3.1-Nemotron-70B | $0.3500/M | $0.4000/M | 128K | $0.3650/M | |
| Qwen2.5 72B Instruct qwen/qwen-2.5-72b-instruct | $0.3600/M | $0.4000/M | 33K | $0.3720/M | |
| Yi-Large-Turbo Yi-Large-Turbo | $0.3800/M | $0.3800/M | 16K | $0.3800/M | |
| GPT-4.1 Mini (batch) openai/gpt-4.1-mini:batch | $0.2000/M | $0.8000/M | 1.0M | $0.3800/M | |
| GPT-5 Mini (batch) openai/gpt-5-mini:batch | $0.1250/M | $1.00/M | 400K | $0.3875/M | |
| Qwen3.5-35B-A3B qwen/qwen3.5-35b-a3b | $0.1400/M | $1.00/M | 262K | $0.3980/M | |
| UnslopNemo 12B thedrummer/unslopnemo-12b | $0.4000/M | $0.4000/M | 1.0M | $0.4000/M | |
| Llama 3.1 70B Instruct meta-llama/llama-3.1-70b-instruct | $0.4000/M | $0.4000/M | 131K | $0.4000/M | |
| Llama 3.1 70B (Hyp) Llama 3.1 70B (Hyp) | $0.4000/M | $0.4000/M | 128K | $0.4000/M | |
| Hermes-3-70B (Hyp) Hermes-3-70B (Hyp) | $0.4000/M | $0.4000/M | 128K | $0.4000/M | |
| Qwen2.5-72B (Hyp) Qwen2.5-72B (Hyp) | $0.4000/M | $0.4000/M | 128K | $0.4000/M | |
| Inception: Mercury 2 inception/mercury-2 | $0.2500/M | $0.7500/M | 128K | $0.4000/M | |
| Mercury 2 inception/mercury-2 | $0.2500/M | $0.7500/M | 128K | $0.4000/M | |
| Moonshot v1 32k Moonshot v1 32k | $0.4000/M | $0.4000/M | 32K | $0.4000/M | |
| Qwen2.5 VL 72B Instruct qwen/qwen2.5-vl-72b-instruct | $0.2500/M | $0.7500/M | 32K | $0.4000/M | |
| GPT-3.5 Turbo (batch) openai/gpt-3.5-turbo:batch | $0.2500/M | $0.7500/M | 16K | $0.4000/M | |
| Granite 20B Multilingual Granite 20B Multilingual | $0.4000/M | $0.4000/M | 8K | $0.4000/M | |
| Qwen3.6 35B A3B qwen/qwen3.6-35b-a3b | $0.1612/M | $0.9653/M | 262K | $0.4024/M | |
| Qwen3 VL 235B A22B Instruct qwen/qwen3-vl-235b-a22b-instruct | $0.2000/M | $0.8800/M | 262K | $0.4040/M | |
| Trinity Large Thinking arcee-ai/trinity-large-thinking | $0.2200/M | $0.8500/M | 262K | $0.4090/M | |
| Venice: Uncensored cognitivecomputations/dolphin-mistral-24b-venice-edition | $0.2000/M | $0.9000/M | 128K | $0.4100/M | |
| Mistral Small 3.1 24B mistralai/mistral-small-3.1-24b-instruct | $0.3510/M | $0.5550/M | 128K | $0.4122/M | |
| Gemini 3.7 Flash (batch) google/gemini-3.7-flash:batch | $0.1875/M | $0.9375/M | 1.0M | $0.4125/M | |
| Qwen Plus 0728 qwen/qwen-plus-2025-07-28 | $0.2600/M | $0.7800/M | 1M | $0.4160/M | |
| Qwen-Plus qwen/qwen-plus | $0.2600/M | $0.7800/M | 1M | $0.4160/M | |
| Qwen3 Coder Flash qwen/qwen3-coder-flash | $0.1950/M | $0.9750/M | 1M | $0.4290/M | |
| MiniMax M2.5 minimax/minimax-m2.5 | $0.1500/M | $1.15/M | 197K | $0.4500/M | |
| ABAB 6.5t ABAB 6.5t | $0.4500/M | $0.4500/M | 8K | $0.4500/M | |
| Hunyuan-Standard Hunyuan-Standard | $0.3500/M | $0.7000/M | 256K | $0.4550/M | |
| Qwen3.6 Flash qwen/qwen3.6-flash | $0.1875/M | $1.13/M | 1M | $0.4687/M | |
| MiniMax-01 minimax/minimax-01 | $0.2000/M | $1.10/M | 1.0M | $0.4700/M | |
| Nex-N2-Pro nex-agi/nex-n2-pro | $0.2500/M | $1.00/M | 262K | $0.4750/M | |
| Gemini 3.5 Flash Lite (batch) google/gemini-3.5-flash-lite:batch | $0.1500/M | $1.25/M | 1.0M | $0.4800/M | |
| Gemini 2.5 Flash (batch) google/gemini-2.5-flash:batch | $0.1500/M | $1.25/M | 1.0M | $0.4800/M | |
| Codestral 2508 mistralai/codestral-2508 | $0.3000/M | $0.9000/M | 256K | $0.4800/M | |
| Z.ai: GLM 4.6V z-ai/glm-4.6v | $0.3000/M | $0.9000/M | 131K | $0.4800/M | |
| Step 3.7 Flash stepfun/step-3.7-flash | $0.2000/M | $1.15/M | 262K | $0.4850/M | |
| DeepSeek V3 deepseek/deepseek-chat | $0.2574/M | $1.03/M | 164K | $0.4888/M | |
| DeepSeek V3.1 Terminus deepseek/deepseek-v3.1-terminus | $0.2700/M | $1.00/M | 164K | $0.4890/M | |
| Qwen3 VL 8B Thinking qwen/qwen3-vl-8b-thinking | $0.1170/M | $1.36/M | 131K | $0.4914/M | |
| GPT-5.6 Luna GPT-5.6 Luna | $0.2000/M | $1.20/M | 1M | $0.5000/M | |
| Mixtral 8x7B (Anyscale) Mixtral 8x7B (Anyscale) | $0.5000/M | $0.5000/M | 32K | $0.5000/M | |
| Mixtral 8x7B (Lepton) Mixtral 8x7B (Lepton) | $0.5000/M | $0.5000/M | 32K | $0.5000/M | |
| DeepSeek-R1-Distill-70B DeepSeek-R1-Distill-70B | $0.3500/M | $0.8800/M | 128K | $0.5090/M | |
| Mixtral 8x7B (Rep) Mixtral 8x7B (Rep) | $0.3000/M | $1.00/M | 32K | $0.5100/M | |
| ReMM SLERP 13B undi95/remm-slerp-l2-13b | $0.4500/M | $0.6500/M | 6K | $0.5100/M | |
| GPT-5.4 Nano openai/gpt-5.4-nano | $0.2000/M | $1.25/M | 400K | $0.5150/M | |
| DeepSeek V3 0324 deepseek/deepseek-chat-v3-0324 | $0.2700/M | $1.12/M | 164K | $0.5250/M | |
| ERNIE 4.5 300B A47B baidu/ernie-4.5-300b-a47b | $0.2800/M | $1.10/M | 123K | $0.5260/M | |
| Muse Glimmer 30B meta/muse-glimmer-30b | $0.3000/M | $1.10/M | 131K | $0.5400/M | |
| Mixtral 8x7B Instruct mistralai/mixtral-8x7b-instruct | $0.5400/M | $0.5400/M | 33K | $0.5400/M | |
| Claude 3 Haiku anthropic/claude-3-haiku | $0.2500/M | $1.25/M | 200K | $0.5500/M | |
| Qwen3 235B A22B Thinking 2507 qwen/qwen3-235b-a22b-thinking-2507 | $0.1495/M | $1.50/M | 131K | $0.5532/M | |
| Perceptron Mk1 perceptron/perceptron-mk1 | $0.1500/M | $1.50/M | 33K | $0.5550/M | |
| Qwen3 VL 30B A3B Thinking qwen/qwen3-vl-30b-a3b-thinking | $0.1300/M | $1.56/M | 131K | $0.5590/M | |
| Jamba Instruct Jamba Instruct | $0.5000/M | $0.7000/M | 256K | $0.5600/M | |
| Granite Code 34B Granite Code 34B | $0.3500/M | $1.05/M | 8K | $0.5600/M | |
| MiMo-V2.5-Pro xiaomi/mimo-v2.5-pro | $0.4350/M | $0.8700/M | 1.1M | $0.5655/M | |
| DeepSeek V4 Pro 0813 deepseek/deepseek-v4-pro-0813 | $0.4350/M | $0.8700/M | 1.0M | $0.5655/M | |
| DeepSeek V4 Pro deepseek/deepseek-v4-pro | $0.4350/M | $0.8700/M | 1.0M | $0.5655/M | |
| LongCat 2.0 meituan/longcat-2.0 | $0.3000/M | $1.20/M | 1.0M | $0.5700/M | |
| MiniMax M3 minimax/minimax-m3 | $0.3000/M | $1.20/M | 1.0M | $0.5700/M | |
| MiniMax M3 (batch) minimax/minimax-m3:batch | $0.3000/M | $1.20/M | 524K | $0.5700/M | |
| KAT-Coder-Pro V2 kwaipilot/kat-coder-pro-v2 | $0.3000/M | $1.20/M | 262K | $0.5700/M | |
| MiniMax M2.1 minimax/minimax-m2.1 | $0.3000/M | $1.20/M | 197K | $0.5700/M | |
| MiniMax M2 minimax/minimax-m2 | $0.3000/M | $1.20/M | 197K | $0.5700/M | |
| MiniMax M2.7 minimax/minimax-m2.7 | $0.3000/M | $1.20/M | 197K | $0.5700/M | |
| MiniMax M2-her minimax/minimax-m2-her | $0.3000/M | $1.20/M | 66K | $0.5700/M | |
| Weaver (alpha) mancer/weaver | $0.5000/M | $0.7500/M | 8K | $0.5750/M | |
| Llama 3 70B Instruct meta-llama/llama-3-70b-instruct | $0.5100/M | $0.7400/M | 8K | $0.5790/M | |
| Grok Code Fast 1 x-ai/grok-code-fast-1 | $0.2000/M | $1.50/M | 256K | $0.5900/M | |
| Phind-CodeLlama-34B (DI) Phind-CodeLlama-34B (DI) | $0.6000/M | $0.6000/M | 16K | $0.6000/M | |
| Yi-34B-Chat (DI) Yi-34B-Chat (DI) | $0.6000/M | $0.6000/M | 4K | $0.6000/M | |
| Qwen3.5-27B qwen/qwen3.5-27b | $0.1950/M | $1.56/M | 262K | $0.6045/M | |
| Qwen3.7 Plus qwen/qwen3.7-plus | $0.3200/M | $1.28/M | 1M | $0.6080/M | |
| WizardLM-2 8x22B microsoft/wizardlm-2-8x22b | $0.6200/M | $0.6200/M | 66K | $0.6200/M | |
| Gemini 3.1 Flash Lite Preview google/gemini-3.1-flash-lite-preview | $0.2500/M | $1.50/M | 1.0M | $0.6250/M | |
| Gemini 3.1 Flash Lite google/gemini-3.1-flash-lite | $0.2500/M | $1.50/M | 1.0M | $0.6250/M | |
| Gemini 3 Flash Preview (batch) google/gemini-3-flash-preview:batch | $0.2500/M | $1.50/M | 1.0M | $0.6250/M | |
| Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) google/gemini-3.1-flash-lite-image | $0.2500/M | $1.50/M | 66K | $0.6250/M | |
| Skyfall 36B V2 thedrummer/skyfall-36b-v2 | $0.5500/M | $0.8000/M | 33K | $0.6250/M | |
| WizardLM-2 8x22B (DI) WizardLM-2 8x22B (DI) | $0.6300/M | $0.6300/M | 64K | $0.6300/M | |
| DeepSeek V3.2 Speciale deepseek/deepseek-v3.2-speciale | $0.4000/M | $1.20/M | 164K | $0.6400/M | |
| Qwen3.5 Plus 2026-02-15 qwen/qwen3.5-plus-02-15 | $0.2600/M | $1.56/M | 1M | $0.6500/M | |
| Llama 3.3 70B (Groq) Llama 3.3 70B (Groq) | $0.5900/M | $0.7900/M | 128K | $0.6500/M | |
| Llama 3.1 70B (Groq) Llama 3.1 70B (Groq) | $0.5900/M | $0.7900/M | 128K | $0.6500/M | |
| Qwen2-57B-A14B Qwen2-57B-A14B | $0.6500/M | $0.6500/M | 64K | $0.6500/M | |
| Gemma 2 27B google/gemma-2-27b-it | $0.6500/M | $0.6500/M | 8K | $0.6500/M | |
| Llama 3 70B Llama 3 70B | $0.5900/M | $0.7900/M | 8K | $0.6500/M | |
| ERNIE 4.5 VL 424B A47B baidu/ernie-4.5-vl-424b-a47b | $0.4200/M | $1.25/M | 123K | $0.6690/M | |
| Thinking Machines: Inkling Small thinkingmachines/inkling-small | $0.4500/M | $1.20/M | 1.0M | $0.6750/M | |
| Llama 3.3 Euryale 70B sao10k/l3.3-euryale-70b | $0.6500/M | $0.7500/M | 131K | $0.6800/M | |
| DeepSeek-R1 (Cerebras) DeepSeek-R1 (Cerebras) | $0.5500/M | $0.9900/M | 64K | $0.6820/M | |
| Qwen3 Coder 480B A35B qwen/qwen3-coder | $0.2200/M | $1.80/M | 262K | $0.6940/M | |
| GLM-4-Long GLM-4-Long | $0.7000/M | $0.7000/M | 1M | $0.7000/M | |
| Nous: Hermes 3 70B Instruct nousresearch/hermes-3-llama-3.1-70b | $0.7000/M | $0.7000/M | 131K | $0.7000/M | |
| Qwen2.5-32B-Instruct Qwen2.5-32B-Instruct | $0.7000/M | $0.7000/M | 128K | $0.7000/M | |
| Qwen2.5-Coder-32B Qwen2.5-Coder-32B | $0.7000/M | $0.7000/M | 128K | $0.7000/M | |
| DeepSeek V4 Flash Vision Exp deepseek/deepseek-v4-flash-vision-exp | $0.4400/M | $1.32/M | 1.0M | $0.7040/M | |
| Llama-3.3-70B (Cerebras) Llama-3.3-70B (Cerebras) | $0.5900/M | $0.9900/M | 128K | $0.7100/M | |
| ERNIE 3.5 8K ERNIE 3.5 8K | $0.5500/M | $1.10/M | 8K | $0.7150/M | |
| Qwen3.5 Plus 2026-04-20 qwen/qwen3.5-plus-20260420 | $0.3000/M | $1.80/M | 1M | $0.7500/M | |
| GPT-4.1 Mini openai/gpt-4.1-mini | $0.4000/M | $1.60/M | 1.0M | $0.7600/M | |
| Qwen2.5 Coder 32B Instruct qwen/qwen-2.5-coder-32b-instruct | $0.6600/M | $1.00/M | 33K | $0.7620/M | |
| GPT-5 Mini openai/gpt-5-mini | $0.2500/M | $2.00/M | 400K | $0.7750/M | |
| GPT-5.1-Codex-Mini openai/gpt-5.1-codex-mini | $0.2500/M | $2.00/M | 400K | $0.7750/M | |
| Seed-2.0-Lite bytedance-seed/seed-2.0-lite | $0.2500/M | $2.00/M | 262K | $0.7750/M | |
| Seed 1.6 bytedance-seed/seed-1.6 | $0.2500/M | $2.00/M | 262K | $0.7750/M | |
| Qwen2.5-72B (Groq) Qwen2.5-72B (Groq) | $0.7900/M | $0.7900/M | 128K | $0.7900/M | |
| Qwen2.5-Coder-32B (Groq) Qwen2.5-Coder-32B (Groq) | $0.7900/M | $0.7900/M | 128K | $0.7900/M | |
| Mistral Large 3 2512 mistralai/mistral-large-2512 | $0.5000/M | $1.50/M | 262K | $0.8000/M | |
| Mistral Large 3 Mistral Large 3 | $0.5000/M | $1.50/M | 256K | $0.8000/M | |
| R1 Distill Llama 70B deepseek/deepseek-r1-distill-llama-70b | $0.8000/M | $0.8000/M | 131K | $0.8000/M | |
| Llama 3.1 405B Llama 3.1 405B | $0.8000/M | $0.8000/M | 128K | $0.8000/M | |
| Aya Expanse 32B Aya Expanse 32B | $0.5000/M | $1.50/M | 128K | $0.8000/M | |
| Llama 3.1 405B (DI) Llama 3.1 405B (DI) | $0.8000/M | $0.8000/M | 128K | $0.8000/M | |
| Llama 3.1 70B (Lepton) Llama 3.1 70B (Lepton) | $0.8000/M | $0.8000/M | 128K | $0.8000/M | |
| Titan Text Premier Titan Text Premier | $0.5000/M | $1.50/M | 32K | $0.8000/M | |
| GPT-3.5 Turbo openai/gpt-3.5-turbo | $0.5000/M | $1.50/M | 16K | $0.8000/M | |
| Phind-CodeLlama-34B (fw) Phind-CodeLlama-34B (fw) | $0.8000/M | $0.8000/M | 16K | $0.8000/M | |
| Aya Expanse 8B Aya Expanse 8B | $0.5000/M | $1.50/M | 8K | $0.8000/M | |
| Nous-Hermes-2-Yi-34B Nous-Hermes-2-Yi-34B | $0.8000/M | $0.8000/M | 4K | $0.8000/M | |
| Yi-34B-Chat Yi-34B-Chat | $0.8000/M | $0.8000/M | 4K | $0.8000/M | |
| Yi-34B Yi-34B | $0.8000/M | $0.8000/M | 4K | $0.8000/M | |
| Z.ai: GLM 4.7 z-ai/glm-4.7 | $0.4000/M | $1.75/M | 205K | $0.8050/M | |
| Qwen3.5-122B-A10B qwen/qwen3.5-122b-a10b | $0.2600/M | $2.08/M | 262K | $0.8060/M | |
| Doubao-Pro-256k Doubao-Pro-256k | $0.5600/M | $1.40/M | 256K | $0.8120/M | |
| Doubao-Pro-128k Doubao-Pro-128k | $0.5600/M | $1.40/M | 128K | $0.8120/M | |
| Qwen3.6 Plus qwen/qwen3.6-plus | $0.3250/M | $1.95/M | 1M | $0.8125/M | |
| DeepSeek-R1 (Groq) DeepSeek-R1 (Groq) | $0.7500/M | $0.9900/M | 128K | $0.8220/M | |
| Gemini 3.7 Flash google/gemini-3.7-flash | $0.3750/M | $1.88/M | 1.0M | $0.8250/M | |
| Gemini 3.6 Flash (batch) google/gemini-3.6-flash:batch | $0.3750/M | $1.88/M | 1.0M | $0.8250/M | |
| Llama 3.1 Euryale 70B v2.2 sao10k/l3.1-euryale-70b | $0.8500/M | $0.8500/M | 131K | $0.8500/M | |
| Qwen3 235B A22B qwen/qwen3-235b-a22b | $0.4550/M | $1.82/M | 131K | $0.8645/M | |
| MoonshotAI: Kimi K2.5 moonshotai/kimi-k2.5 | $0.3750/M | $2.02/M | 262K | $0.8700/M | |
| MoonshotAI: Kimi K2 0905 moonshotai/kimi-k2-0905 | $0.4000/M | $2.00/M | 262K | $0.8800/M | |
| Devstral Medium mistralai/devstral-medium | $0.4000/M | $2.00/M | 131K | $0.8800/M | |
| Mistral Medium 3 mistralai/mistral-medium-3 | $0.4000/M | $2.00/M | 131K | $0.8800/M | |
| Mistral Medium 3.1 mistralai/mistral-medium-3.1 | $0.4000/M | $2.00/M | 131K | $0.8800/M | |
| Reka Edge Reka Edge | $0.4000/M | $2.00/M | 128K | $0.8800/M | |
| Llama 3.2 90B Vision Llama 3.2 90B Vision | $0.8800/M | $0.8800/M | 128K | $0.8800/M | |
| Llama 3.3 70B (Together) Llama 3.3 70B (Together) | $0.8800/M | $0.8800/M | 128K | $0.8800/M | |
| Virtuoso Large arcee-ai/virtuoso-large | $0.7500/M | $1.20/M | 131K | $0.8850/M | |
| Qwen2-72B-Instruct Qwen2-72B-Instruct | $0.9000/M | $0.9000/M | 128K | $0.9000/M | |
| Mixtral 8x22B Mixtral 8x22B | $0.9000/M | $0.9000/M | 64K | $0.9000/M | |
| Qwen2.5-72B (fw) Qwen2.5-72B (fw) | $0.9000/M | $0.9000/M | 32K | $0.9000/M | |
| CodeLlama 70B Instruct CodeLlama 70B Instruct | $0.9000/M | $0.9000/M | 16K | $0.9000/M | |
| Llama 2 70B Chat Llama 2 70B Chat | $0.9000/M | $0.9000/M | 4K | $0.9000/M | |
| Yi-34B (fw) Yi-34B (fw) | $0.9000/M | $0.9000/M | 4K | $0.9000/M | |
| Falcon 40B (Together) Falcon 40B (Together) | $0.9000/M | $0.9000/M | 2K | $0.9000/M | |
| Falcon 40B Instruct Falcon 40B Instruct | $0.9000/M | $0.9000/M | 2K | $0.9000/M | |
| Hunyuan-Standard-256K Hunyuan-Standard-256K | $0.7000/M | $1.40/M | 256K | $0.9100/M | |
| AionLabs: Aion-3.0-Mini aion-labs/aion-3.0-mini | $0.7000/M | $1.40/M | 131K | $0.9100/M | |
| Morph V3 Fast morph/morph-v3-fast | $0.8000/M | $1.20/M | 82K | $0.9200/M | |
| Qwen3.6 27B qwen/qwen3.6-27b | $0.2890/M | $2.40/M | 256K | $0.9223/M | |
| GPT-5.4 Mini (batch) openai/gpt-5.4-mini:batch | $0.3750/M | $2.25/M | 400K | $0.9375/M | |
| MiniMax M1 minimax/minimax-m1 | $0.4000/M | $2.20/M | 1M | $0.9400/M | |
| Z.ai: GLM 4.6 z-ai/glm-4.6 | $0.5000/M | $2.00/M | 205K | $0.9500/M | |
| Palmyra X 004 Palmyra X 004 | $0.5000/M | $2.00/M | 128K | $0.9500/M | |
| Gemini 2.5 Flash google/gemini-2.5-flash | $0.3000/M | $2.50/M | 1.0M | $0.9600/M | |
| Gemini 3.5 Flash-Lite google/gemini-3.5-flash-lite | $0.3000/M | $2.50/M | 1.0M | $0.9600/M | |
| Gemini 3.5 Flash Lite google/gemini-3.5-flash-lite | $0.3000/M | $2.50/M | 1.0M | $0.9600/M | |
| Nova 2 Lite amazon/nova-2-lite-v1 | $0.3000/M | $2.50/M | 1M | $0.9600/M | |
| Z.ai: GLM 4.5V z-ai/glm-4.5v | $0.6000/M | $1.80/M | 66K | $0.9600/M | |
| Nano Banana (Gemini 2.5 Flash Image) google/gemini-2.5-flash-image | $0.3000/M | $2.50/M | 33K | $0.9600/M | |
| Qwen3 VL 235B A22B Thinking qwen/qwen3-vl-235b-a22b-thinking | $0.2600/M | $2.60/M | 131K | $0.9620/M | |
| Devstral 2 2512 mistralai/devstral-2512 | $0.4400/M | $2.20/M | 262K | $0.9680/M | |
| Relace Apply 3 relace/relace-apply-3 | $0.8500/M | $1.25/M | 256K | $0.9700/M | |
| Qwen VL Max qwen/qwen-vl-max | $0.5200/M | $2.08/M | 131K | $0.9880/M | |
| Llama-3.1-405B (Cerebras) Llama-3.1-405B (Cerebras) | $0.9900/M | $0.9900/M | 128K | $0.9900/M | |
| R1 0528 deepseek/deepseek-r1-0528 | $0.5000/M | $2.15/M | 164K | $0.9950/M | |
| Z.ai: GLM 5 z-ai/glm-5 | $0.6000/M | $1.92/M | 205K | $0.9960/M | |
| Nous: Hermes 3 405B Instruct nousresearch/hermes-3-llama-3.1-405b | $1.00/M | $1.00/M | 131K | $1.00/M | |
| Llama 3.1 70B (Anyscale) Llama 3.1 70B (Anyscale) | $1.00/M | $1.00/M | 128K | $1.00/M | |
| Sonar perplexity/sonar | $1.00/M | $1.00/M | 127K | $1.00/M | |
| CodeLlama 70B (Anyscale) CodeLlama 70B (Anyscale) | $1.00/M | $1.00/M | 16K | $1.00/M | |
| DeepSeek V4 Pro 0423 deepseek/deepseek-v4-pro | $0.7903/M | $1.58/M | 1.0M | $1.03/M | |
| AionLabs: Aion-2.0 aion-labs/aion-2.0 | $0.8000/M | $1.60/M | 131K | $1.04/M | |
| AionLabs: Aion-RP 1.0 (8B) aion-labs/aion-rp-llama-3.1-8b | $0.8000/M | $1.60/M | 33K | $1.04/M | |
| DeepSeek-R1 DeepSeek-R1 | $0.5500/M | $2.19/M | 64K | $1.04/M | |
| DeepSeek-R1-0528 DeepSeek-R1-0528 | $0.5500/M | $2.19/M | 64K | $1.04/M | |
| DeepSeek-R1 (DI) DeepSeek-R1 (DI) | $0.5500/M | $2.19/M | 64K | $1.04/M | |
| o4 Mini High (batch) openai/o4-mini-high:batch | $0.5500/M | $2.20/M | 200K | $1.04/M | |
| o4 Mini (batch) openai/o4-mini:batch | $0.5500/M | $2.20/M | 200K | $1.04/M | |
| o3 Mini High (batch) openai/o3-mini-high:batch | $0.5500/M | $2.20/M | 200K | $1.04/M | |
| o3 Mini (batch) openai/o3-mini:batch | $0.5500/M | $2.20/M | 200K | $1.04/M | |
| Z.ai: GLM 4.5 z-ai/glm-4.5 | $0.6000/M | $2.20/M | 131K | $1.08/M | |
| MoonshotAI: Kimi K2 0711 moonshotai/kimi-k2 | $0.5700/M | $2.30/M | 131K | $1.09/M | |
| Kimi K2 0711 moonshotai/kimi-k2 | $0.5700/M | $2.30/M | 131K | $1.09/M | |
| Nemotron 3 Ultra nvidia/nemotron-3-ultra-550b-a55b | $0.5000/M | $2.50/M | 1M | $1.10/M | |
| Seed 2.1 Turbo bytedance-seed/seed-2-1-turbo | $0.5000/M | $2.50/M | 262K | $1.10/M | |
| Claude Haiku 4.5 (batch) anthropic/claude-haiku-4.5:batch | $0.5000/M | $2.50/M | 200K | $1.10/M | |
| MoonshotAI: Kimi K2.6 moonshotai/kimi-k2.6 | $0.5605/M | $2.36/M | 262K | $1.10/M | |
| QwQ-32B QwQ-32B | $0.6000/M | $2.40/M | 131K | $1.14/M | |
| GPT Audio Mini openai/gpt-audio-mini | $0.6000/M | $2.40/M | 128K | $1.14/M | |
| MoonshotAI: Kimi K2 Thinking moonshotai/kimi-k2-thinking | $0.6000/M | $2.50/M | 262K | $1.17/M | |
| Kimi K2 Thinking moonshotai/kimi-k2-thinking | $0.6000/M | $2.50/M | 262K | $1.17/M | |
| Kimi K2 0905 moonshotai/kimi-k2-0905 | $0.6000/M | $2.50/M | 262K | $1.17/M | |
| DBRX Instruct DBRX Instruct | $0.7500/M | $2.25/M | 32K | $1.20/M | |
| DBRX Base DBRX Base | $0.7500/M | $2.25/M | 32K | $1.20/M | |
| Morph V3 Large morph/morph-v3-large | $0.9000/M | $1.90/M | 262K | $1.20/M | |
| Llama 3.1 Nemotron 70B Instruct nvidia/llama-3.1-nemotron-70b-instruct | $1.20/M | $1.20/M | 131K | $1.20/M | |
| Qwen2.5-72B-Instruct Qwen2.5-72B-Instruct | $1.20/M | $1.20/M | 128K | $1.20/M | |
| Mixtral 8x22B (Together) Mixtral 8x22B (Together) | $1.20/M | $1.20/M | 64K | $1.20/M | |
| WizardLM-2 8x22B WizardLM-2 8x22B | $1.20/M | $1.20/M | 64K | $1.20/M | |
| Qwen2.5-72B (Together) Qwen2.5-72B (Together) | $1.20/M | $1.20/M | 32K | $1.20/M | |
| DBRX Instruct (Together) DBRX Instruct (Together) | $1.20/M | $1.20/M | 32K | $1.20/M | |
| Qwen3.5 397B A17B qwen/qwen3.5-397b-a17b | $0.4500/M | $3.00/M | 262K | $1.21/M | |
| R1 deepseek/deepseek-r1 | $0.7000/M | $2.50/M | 64K | $1.24/M | |
| Gemini 3 Flash Preview google/gemini-3-flash-preview | $0.5000/M | $3.00/M | 1.0M | $1.25/M | |
| Gemini 3 Flash Gemini 3 Flash | $0.5000/M | $3.00/M | 1M | $1.25/M | |
| Seed-2.0-Code bytedance-seed/seed-2.0-code | $0.5000/M | $3.00/M | 262K | $1.25/M | |
| Nano Banana 2 (Gemini 3.1 Flash Image) google/gemini-3.1-flash-image | $0.5000/M | $3.00/M | 131K | $1.25/M | |
| Nano Banana 2 (Gemini 3.1 Flash Image Preview) google/gemini-3.1-flash-image-preview | $0.5000/M | $3.00/M | 66K | $1.25/M | |
| Qwen3.8 27B qwen/qwen3.8-27b | $0.4500/M | $3.20/M | 262K | $1.27/M | |
| Grok Build 0.1 x-ai/grok-build-0.1 | $1.00/M | $2.00/M | 256K | $1.30/M | |
| SpaceXAI: Grok Build 0.1 x-ai/grok-build-0.1 | $1.00/M | $2.00/M | 256K | $1.30/M | |
| GPT-3.5 Turbo (older v0613) openai/gpt-3.5-turbo-0613 | $1.00/M | $2.00/M | 4K | $1.30/M | |
| Qianfan-OCR-Fast baidu/qianfan-ocr-fast | $0.6800/M | $2.81/M | 66K | $1.32/M | |
| Kimi K2.5 moonshotai/kimi-k2.5 | $0.6000/M | $3.00/M | 262K | $1.32/M | |
| MiniMax-Text-01 MiniMax-Text-01 | $0.7000/M | $2.80/M | 1M | $1.33/M | |
| KAT-Coder-Pro V2.5 kwaipilot/kat-coder-pro-v2.5 | $0.7400/M | $2.96/M | 256K | $1.41/M | |
| Qwen3 Coder Plus qwen/qwen3-coder-plus | $0.6500/M | $3.25/M | 1M | $1.43/M | |
| MoonshotAI: Kimi K2.7 Code moonshotai/kimi-k2.7-code | $0.6700/M | $3.40/M | 262K | $1.49/M | |
| Kimi K2.7 Code moonshotai/kimi-k2.7-code | $0.6700/M | $3.40/M | 262K | $1.49/M | |
| Nemotron 3 Ultra (batch) nvidia/nemotron-3-ultra-550b-a55b:batch | $0.6000/M | $3.60/M | 512K | $1.50/M | |
| Nova 2 Pro Nova 2 Pro | $0.8000/M | $3.20/M | 1M | $1.52/M | |
| Nova Pro 1.0 amazon/nova-pro-v1 | $0.8000/M | $3.20/M | 300K | $1.52/M | |
| Nova Pro Nova Pro | $0.8000/M | $3.20/M | 300K | $1.52/M | |
| Relace Search relace/relace-search | $1.00/M | $3.00/M | 256K | $1.60/M | |
| Nous: Hermes 4 405B nousresearch/hermes-4-405b | $1.00/M | $3.00/M | 131K | $1.60/M | |
| Llama-3.1-Nemotron-Ultra-253B Llama-3.1-Nemotron-Ultra-253B | $1.60/M | $1.60/M | 128K | $1.60/M | |
| Llama 3.1 70B (DB) Llama 3.1 70B (DB) | $1.00/M | $3.00/M | 128K | $1.60/M | |
| Palmyra X 32k Palmyra X 32k | $1.00/M | $3.00/M | 32K | $1.60/M | |
| Palmyra Med Palmyra Med | $1.00/M | $3.00/M | 32K | $1.60/M | |
| Palmyra Fin Palmyra Fin | $1.00/M | $3.00/M | 32K | $1.60/M | |
| Grok 4.20 x-ai/grok-4.20 | $1.25/M | $2.50/M | 2M | $1.63/M | |
| SpaceXAI: Grok 4.20 Multi-Agent x-ai/grok-4.20-multi-agent | $1.25/M | $2.50/M | 2M | $1.63/M | |
| SpaceXAI: Grok 4.20 x-ai/grok-4.20 | $1.25/M | $2.50/M | 2M | $1.63/M | |
| Grok 4.3 x-ai/grok-4.3 | $1.25/M | $2.50/M | 1M | $1.63/M | |
| SpaceXAI: Grok 4.3 x-ai/grok-4.3 | $1.25/M | $2.50/M | 1M | $1.63/M | |
| GPT-3.5 Turbo Instruct openai/gpt-3.5-turbo-instruct | $1.50/M | $2.00/M | 4K | $1.65/M | |
| Qwen3 Max Thinking qwen/qwen3-max-thinking | $0.7800/M | $3.90/M | 262K | $1.72/M | |
| Qwen3 Max qwen/qwen3-max | $0.7800/M | $3.90/M | 262K | $1.72/M | |
| Claude Haiku 3.5 Claude Haiku 3.5 | $0.8000/M | $4.00/M | 200K | $1.76/M | |
| Claude 3.5 Haiku anthropic/claude-3.5-haiku | $0.8000/M | $4.00/M | 200K | $1.76/M | |
| Reka Flash 21B Reka Flash 21B | $0.8000/M | $4.00/M | 128K | $1.76/M | |
| Moonshot v1 128k Moonshot v1 128k | $1.80/M | $1.80/M | 128K | $1.80/M | |
| Sakana Namazu sakana/sakana-namazu | $0.9500/M | $4.00/M | 262K | $1.86/M | |
| Kimi K2.7 Code (batch) moonshotai/kimi-k2.7-code:batch | $0.9500/M | $4.00/M | 262K | $1.86/M | |
| Kimi K2.6 moonshotai/kimi-k2.6 | $0.9500/M | $4.00/M | 262K | $1.86/M | |
| Gemini 3.5 Flash (batch) google/gemini-3.5-flash:batch | $0.7500/M | $4.50/M | 1.0M | $1.87/M | |
| GPT-5.4 Mini openai/gpt-5.4-mini | $0.7500/M | $4.50/M | 400K | $1.87/M | |
| Thinking Machines: Inkling thinkingmachines/inkling | $0.9500/M | $4.05/M | 1.0M | $1.88/M | |
| GPT-4.1 (batch) openai/gpt-4.1:batch | $1.00/M | $4.00/M | 1.0M | $1.90/M | |
| o3 (batch) openai/o3:batch | $1.00/M | $4.00/M | 200K | $1.90/M | |
| Thinking Machines: Inkling (batch) thinkingmachines/inkling:batch | $1.00/M | $4.05/M | 524K | $1.91/M | |
| Gemini 2.5 Pro (batch) google/gemini-2.5-pro:batch | $0.6250/M | $5.00/M | 1.0M | $1.94/M | |
| GPT-5.1 (batch) openai/gpt-5.1:batch | $0.6250/M | $5.00/M | 400K | $1.94/M | |
| GPT-5 Codex (batch) openai/gpt-5-codex:batch | $0.6250/M | $5.00/M | 400K | $1.94/M | |
| GPT-5 (batch) openai/gpt-5:batch | $0.6250/M | $5.00/M | 400K | $1.94/M | |
| Z.ai: GLM 5.2 z-ai/glm-5.2 | $1.19/M | $3.74/M | 1.0M | $1.96/M | |
| Qwen-Max qwen/qwen-max | $1.04/M | $4.16/M | 33K | $1.98/M | |
| Qwen2-VL-72B Qwen2-VL-72B | $2.00/M | $2.00/M | 32K | $2.00/M | |
| Rerank v3.5 Rerank v3.5 | $2.00/M | — | — | $2.00/M | |
| Rerank v3 Multilingual Rerank v3 Multilingual | $2.00/M | — | — | $2.00/M | |
| Z.ai: GLM 5V Turbo z-ai/glm-5v-turbo | $1.20/M | $4.00/M | 203K | $2.04/M | |
| Z.ai: GLM 5 Turbo z-ai/glm-5-turbo | $1.20/M | $4.00/M | 203K | $2.04/M | |
| Z.ai: GLM 5.1 z-ai/glm-5.1 | $1.26/M | $3.96/M | 205K | $2.07/M | |
| o3 Mini openai/o3-mini | $1.10/M | $4.40/M | 200K | $2.09/M | |
| o4 Mini openai/o4-mini | $1.10/M | $4.40/M | 200K | $2.09/M | |
| o3 Mini High openai/o3-mini-high | $1.10/M | $4.40/M | 200K | $2.09/M | |
| o4 Mini High openai/o4-mini-high | $1.10/M | $4.40/M | 200K | $2.09/M | |
| Palmyra Vision Palmyra Vision | $1.50/M | $3.50/M | 32K | $2.10/M | |
| Muse Spark 1.2 meta/muse-spark-1.2 | $1.25/M | $4.25/M | 1.0M | $2.15/M | |
| Muse Spark 1.1 meta/muse-spark-1.1 | $1.25/M | $4.25/M | 1.0M | $2.15/M | |
| GPT-5.6 Sol Pro (batch) openai/gpt-5.6-sol-pro:batch | $1.00/M | $5.00/M | 1.1M | $2.20/M | |
| GPT-5.6 Sol (batch) openai/gpt-5.6-sol:batch | $1.00/M | $5.00/M | 1.1M | $2.20/M | |
| Claude Sonnet 5 (batch) anthropic/claude-sonnet-5:batch | $1.00/M | $5.00/M | 1M | $2.20/M | |
| Claude Haiku 4.5 anthropic/claude-haiku-4.5 | $1.00/M | $5.00/M | 200K | $2.20/M | |
| Sonar Reasoning Sonar Reasoning | $1.00/M | $5.00/M | 127K | $2.20/M | |
| ABAB 6.5s ABAB 6.5s | $1.00/M | $5.00/M | 8K | $2.20/M | |
| Palmyra X5 writer/palmyra-x5 | $0.6000/M | $6.00/M | 1.0M | $2.22/M | |
| Z.ai: GLM 5.3 z-ai/glm-5.3 | $1.40/M | $4.40/M | 1.0M | $2.30/M | |
| Z.ai: GLM 5.2 (batch) z-ai/glm-5.2:batch | $1.40/M | $4.40/M | 1.0M | $2.30/M | |
| GPT-5 Image Mini openai/gpt-5-image-mini | $2.50/M | $2.00/M | 400K | $2.35/M | |
| Gemini 1.5 Pro Gemini 1.5 Pro | $1.25/M | $5.00/M | 2M | $2.38/M | |
| GPT-4o (batch) openai/gpt-4o:batch | $1.25/M | $5.00/M | 128K | $2.38/M | |
| GPT-5.6 Luna Pro openai/gpt-5.6-luna-pro | $1.00/M | $6.00/M | 1.1M | $2.50/M | |
| GPT-5.6 Terra Pro (batch) openai/gpt-5.6-terra-pro:batch | $1.00/M | $6.00/M | 1.1M | $2.50/M | |
| GPT-5.6 Terra (batch) openai/gpt-5.6-terra:batch | $1.00/M | $6.00/M | 1.1M | $2.50/M | |
| Gemini 3.1 Pro Preview (batch) google/gemini-3.1-pro-preview:batch | $1.00/M | $6.00/M | 1.0M | $2.50/M | |
| Qwen3.6 Max Preview qwen/qwen3.6-max-preview | $1.03/M | $6.16/M | 262K | $2.57/M | |
| GPT-5.2 (batch) openai/gpt-5.2:batch | $0.8750/M | $7.00/M | 400K | $2.71/M | |
| Solar Pro Solar Pro | $1.50/M | $6.00/M | 32K | $2.85/M | |
| Llama 3.1 405B (Groq) Llama 3.1 405B (Groq) | $2.99/M | $2.99/M | 128K | $2.99/M | |
| Llama 3.1 405B (fw) Llama 3.1 405B (fw) | $3.00/M | $3.00/M | 131K | $3.00/M | |
| Yi-Large Yi-Large | $3.00/M | $3.00/M | 32K | $3.00/M | |
| Qwen2.5-Max Qwen2.5-Max | $1.60/M | $6.40/M | 32K | $3.04/M | |
| GPT-5.4 (batch) openai/gpt-5.4:batch | $1.25/M | $7.50/M | 1.1M | $3.13/M | |
| Grok 4.20 Multi-Agent x-ai/grok-4.20-multi-agent | $2.00/M | $6.00/M | 2M | $3.20/M | |
| Qwen3.8 Max qwen/qwen3.8-max | $2.00/M | $6.00/M | 1M | $3.20/M | |
| SpaceXAI: Grok 4.6 x-ai/grok-4.6 | $2.00/M | $6.00/M | 500K | $3.20/M | |
| Grok 4.5 x-ai/grok-4.5 | $2.00/M | $6.00/M | 500K | $3.20/M | |
| SpaceXAI: Grok 4.5 x-ai/grok-4.5 | $2.00/M | $6.00/M | 500K | $3.20/M | |
| Qwen3.8 2.4T A95B qwen/qwen3.8-2.4t-a95b | $2.00/M | $6.00/M | 262K | $3.20/M | |
| Mistral Large 2407 mistralai/mistral-large-2407 | $2.00/M | $6.00/M | 131K | $3.20/M | |
| Mistral Large 2411 mistralai/mistral-large-2411 | $2.00/M | $6.00/M | 131K | $3.20/M | |
| Pixtral Large 2411 mistralai/pixtral-large-2411 | $2.00/M | $6.00/M | 131K | $3.20/M | |
| Mistral Large 2 Mistral Large 2 | $2.00/M | $6.00/M | 128K | $3.20/M | |
| Pixtral Large Pixtral Large | $2.00/M | $6.00/M | 128K | $3.20/M | |
| Mistral-Large-2 (NIM) Mistral-Large-2 (NIM) | $2.00/M | $6.00/M | 128K | $3.20/M | |
| Mistral Large mistralai/mistral-large | $2.00/M | $6.00/M | 128K | $3.20/M | |
| Mixtral 8x22B Instruct mistralai/mixtral-8x22b-instruct | $2.00/M | $6.00/M | 66K | $3.20/M | |
| Gemini 3.6 Flash google/gemini-3.6-flash | $1.50/M | $7.50/M | 1.0M | $3.30/M | |
| Claude Sonnet 4.6 (batch) anthropic/claude-sonnet-4.6:batch | $1.50/M | $7.50/M | 1M | $3.30/M | |
| Claude Sonnet 4.5 (batch) anthropic/claude-sonnet-4.5:batch | $1.50/M | $7.50/M | 1M | $3.30/M | |
| Mistral Medium 3.5 mistralai/mistral-medium-3-5 | $1.50/M | $7.50/M | 262K | $3.30/M | |
| GPT-3.5 Turbo 16k openai/gpt-3.5-turbo-16k | $3.00/M | $4.00/M | 16K | $3.30/M | |
| Llama 3.1 405B (Together) Llama 3.1 405B (Together) | $3.50/M | $3.50/M | 128K | $3.50/M | |
| Yi-VL-34B Yi-VL-34B | $3.50/M | $3.50/M | 4K | $3.50/M | |
| Magnum v4 72B anthracite-org/magnum-v4-72b | $3.00/M | $5.00/M | 33K | $3.60/M | |
| Gemini 3.5 Flash google/gemini-3.5-flash | $1.50/M | $9.00/M | 1.0M | $3.75/M | |
| GPT-4.1 openai/gpt-4.1 | $2.00/M | $8.00/M | 1.0M | $3.80/M | |
| Jamba 1.5 Large Jamba 1.5 Large | $2.00/M | $8.00/M | 256K | $3.80/M | |
| Jamba Large 1.7 ai21/jamba-large-1.7 | $2.00/M | $8.00/M | 256K | $3.80/M | |
| o3 openai/o3 | $2.00/M | $8.00/M | 200K | $3.80/M | |
| o4 Mini Deep Research openai/o4-mini-deep-research | $2.00/M | $8.00/M | 200K | $3.80/M | |
| Sonar Reasoning Pro perplexity/sonar-reasoning-pro | $2.00/M | $8.00/M | 128K | $3.80/M | |
| Sonar Deep Research perplexity/sonar-deep-research | $2.00/M | $8.00/M | 128K | $3.80/M | |
| Gemini 2.5 Pro google/gemini-2.5-pro | $1.25/M | $10.00/M | 1.0M | $3.88/M | |
| Gemini 2.5 Pro Preview 06-05 google/gemini-2.5-pro-preview | $1.25/M | $10.00/M | 1.0M | $3.88/M | |
| Gemini 2.5 Pro Preview 05-06 google/gemini-2.5-pro-preview-05-06 | $1.25/M | $10.00/M | 1.0M | $3.88/M | |
| GPT-5 openai/gpt-5 | $1.25/M | $10.00/M | 400K | $3.88/M | |
| GPT-5 Codex openai/gpt-5-codex | $1.25/M | $10.00/M | 400K | $3.88/M | |
| GPT-5.1 openai/gpt-5.1 | $1.25/M | $10.00/M | 400K | $3.88/M | |
| GPT-5.1-Codex-Max openai/gpt-5.1-codex-max | $1.25/M | $10.00/M | 400K | $3.88/M | |
| GPT-5.1-Codex openai/gpt-5.1-codex | $1.25/M | $10.00/M | 400K | $3.88/M | |
| GPT-5 Chat openai/gpt-5-chat | $1.25/M | $10.00/M | 128K | $3.88/M | |
| GPT-5.1 Chat openai/gpt-5.1-chat | $1.25/M | $10.00/M | 128K | $3.88/M | |
| AionLabs: Aion-3.0 aion-labs/aion-3.0 | $3.00/M | $6.00/M | 131K | $3.90/M | |
| Qwen3.7 Max qwen/qwen3.7-max | $2.50/M | $7.50/M | 1M | $4.00/M | |
| Llama 3.1 405B (Hyp) Llama 3.1 405B (Hyp) | $4.00/M | $4.00/M | 128K | $4.00/M | |
| DeepSeek-R1 (Together) DeepSeek-R1 (Together) | $3.00/M | $7.00/M | 64K | $4.20/M | |
| Nemotron-4-340B Nemotron-4-340B | $4.20/M | $4.20/M | 4K | $4.20/M | |
| Claude Sonnet 5 Claude Sonnet 5 | $2.00/M | $10.00/M | 1M | $4.40/M | |
| Grok 2 Grok 2 | $2.00/M | $10.00/M | 131K | $4.40/M | |
| Grok 2 Vision Grok 2 Vision | $2.00/M | $10.00/M | 32K | $4.40/M | |
| DeepSeek-R1 (fw) DeepSeek-R1 (fw) | $3.00/M | $8.00/M | 64K | $4.50/M | |
| Hunyuan-Turbo Hunyuan-Turbo | $3.50/M | $7.00/M | 256K | $4.55/M | |
| Command A cohere/command-a | $2.50/M | $10.00/M | 256K | $4.75/M | |
| Command A+ Command A+ | $2.50/M | $10.00/M | 256K | $4.75/M | |
| GPT-4o openai/gpt-4o | $2.50/M | $10.00/M | 128K | $4.75/M | |
| Command R+ (08-2024) cohere/command-r-plus-08-2024 | $2.50/M | $10.00/M | 128K | $4.75/M | |
| Kimi k1.5 Kimi k1.5 | $2.50/M | $10.00/M | 128K | $4.75/M | |
| GPT-4o Audio openai/gpt-4o-audio-preview | $2.50/M | $10.00/M | 128K | $4.75/M | |
| GPT-4o Search Preview openai/gpt-4o-search-preview | $2.50/M | $10.00/M | 128K | $4.75/M | |
| GPT Audio openai/gpt-audio | $2.50/M | $10.00/M | 128K | $4.75/M | |
| Pi (Inflection-2.5) Pi (Inflection-2.5) | $2.50/M | $10.00/M | 32K | $4.75/M | |
| Inflection 3 Pi inflection/inflection-3-pi | $2.50/M | $10.00/M | 8K | $4.75/M | |
| Inflection 3 Productivity inflection/inflection-3-productivity | $2.50/M | $10.00/M | 8K | $4.75/M | |
| GPT-5.6 Terra Pro openai/gpt-5.6-terra-pro | $2.00/M | $12.00/M | 1.1M | $5.00/M | |
| Gemini 3.1 Pro Preview google/gemini-3.1-pro-preview | $2.00/M | $12.00/M | 1.0M | $5.00/M | |
| Gemini 3.1 Pro Preview Custom Tools google/gemini-3.1-pro-preview-customtools | $2.00/M | $12.00/M | 1.0M | $5.00/M | |
| GPT-5.6 Terra GPT-5.6 Terra | $2.00/M | $12.00/M | 1M | $5.00/M | |
| Gemini 3.5 Pro Gemini 3.5 Pro | $2.00/M | $12.00/M | 1M | $5.00/M | |
| Gemini 3.1 Pro Gemini 3.1 Pro | $2.00/M | $12.00/M | 1M | $5.00/M | |
| Gemini 3 Pro Gemini 3 Pro | $2.00/M | $12.00/M | 1M | $5.00/M | |
| Sonar Huge Sonar Huge | $5.00/M | $5.00/M | 127K | $5.00/M | |
| Nano Banana Pro (Gemini 3 Pro Image) google/gemini-3-pro-image | $2.00/M | $12.00/M | 66K | $5.00/M | |
| Nano Banana Pro (Gemini 3 Pro Image Preview) google/gemini-3-pro-image-preview | $2.00/M | $12.00/M | 66K | $5.00/M | |
| GPT-5.2 openai/gpt-5.2 | $1.75/M | $14.00/M | 400K | $5.42/M | |
| GPT-5.2-Codex openai/gpt-5.2-codex | $1.75/M | $14.00/M | 400K | $5.42/M | |
| GPT-5.3-Codex openai/gpt-5.3-codex | $1.75/M | $14.00/M | 400K | $5.42/M | |
| GPT-5.3 Chat openai/gpt-5.3-chat | $1.75/M | $14.00/M | 128K | $5.42/M | |
| GPT-5.2 Chat openai/gpt-5.2-chat | $1.75/M | $14.00/M | 128K | $5.42/M | |
| Falcon 180B Falcon 180B | $3.50/M | $10.00/M | 4K | $5.45/M | |
| Nova Premier 1.0 amazon/nova-premier-v1 | $2.50/M | $12.50/M | 1M | $5.50/M | |
| Claude Opus 5 (batch) anthropic/claude-opus-5:batch | $2.50/M | $12.50/M | 1M | $5.50/M | |
| Claude Opus 4.8 (batch) anthropic/claude-opus-4.8:batch | $2.50/M | $12.50/M | 1M | $5.50/M | |
| Claude Opus 4.7 (batch) anthropic/claude-opus-4.7:batch | $2.50/M | $12.50/M | 1M | $5.50/M | |
| Claude Opus 4.6 (batch) anthropic/claude-opus-4.6:batch | $2.50/M | $12.50/M | 1M | $5.50/M | |
| Nova Premier Nova Premier | $2.50/M | $12.50/M | 300K | $5.50/M | |
| Claude Opus 4.5 (batch) anthropic/claude-opus-4.5:batch | $2.50/M | $12.50/M | 200K | $5.50/M | |
| Hunyuan-Vision Hunyuan-Vision | $5.50/M | $5.50/M | 4K | $5.50/M | |
| ERNIE 4.0 Turbo 8K ERNIE 4.0 Turbo 8K | $3.50/M | $10.50/M | 8K | $5.60/M | |
| GPT-5.4 openai/gpt-5.4 | $2.50/M | $15.00/M | 1.1M | $6.25/M | |
| GPT-5.5 (batch) openai/gpt-5.5:batch | $2.50/M | $15.00/M | 1.1M | $6.25/M | |
| MoonshotAI: Kimi K3 moonshotai/kimi-k3 | $3.00/M | $15.00/M | 1.0M | $6.60/M | |
| Kimi K3 moonshotai/kimi-k3 | $3.00/M | $15.00/M | 1.0M | $6.60/M | |
| Claude Sonnet 4.5 anthropic/claude-sonnet-4.5 | $3.00/M | $15.00/M | 1M | $6.60/M | |
| Claude Sonnet 4.6 anthropic/claude-sonnet-4.6 | $3.00/M | $15.00/M | 1M | $6.60/M | |
| Claude Sonnet 4 anthropic/claude-sonnet-4 | $3.00/M | $15.00/M | 1M | $6.60/M | |
| Grok 4 x-ai/grok-4 | $3.00/M | $15.00/M | 256K | $6.60/M | |
| Claude Sonnet 3.7 Claude Sonnet 3.7 | $3.00/M | $15.00/M | 200K | $6.60/M | |
| Claude Sonnet 3.5 Claude Sonnet 3.5 | $3.00/M | $15.00/M | 200K | $6.60/M | |
| Sonar Pro perplexity/sonar-pro | $3.00/M | $15.00/M | 200K | $6.60/M | |
| Sonar Pro Search perplexity/sonar-pro-search | $3.00/M | $15.00/M | 200K | $6.60/M | |
| Claude 3.7 Sonnet anthropic/claude-3.7-sonnet | $3.00/M | $15.00/M | 200K | $6.60/M | |
| Grok 3 x-ai/grok-3 | $3.00/M | $15.00/M | 131K | $6.60/M | |
| Grok 3 Beta x-ai/grok-3-beta | $3.00/M | $15.00/M | 131K | $6.60/M | |
| Reka Core Reka Core | $3.00/M | $15.00/M | 128K | $6.60/M | |
| ERNIE 4.0 8K ERNIE 4.0 8K | $4.20/M | $12.60/M | 8K | $6.72/M | |
| GLM-4-Plus GLM-4-Plus | $7.00/M | $7.00/M | 128K | $7.00/M | |
| GLM-4V GLM-4V | $7.00/M | $7.00/M | 2K | $7.00/M | |
| GPT-4 Turbo (batch) openai/gpt-4-turbo:batch | $5.00/M | $15.00/M | 128K | $8.00/M | |
| Llama 3.1 405B (Rep) Llama 3.1 405B (Rep) | $9.50/M | $9.50/M | 128K | $9.50/M | |
| GPT-5 Image openai/gpt-5-image | $10.00/M | $10.00/M | 400K | $10.00/M | |
| GPT-5.4 Image 2 openai/gpt-5.4-image-2 | $8.00/M | $15.00/M | 272K | $10.10/M | |
| Claude Opus 4.6 anthropic/claude-opus-4.6 | $5.00/M | $25.00/M | 1M | $11.00/M | |
| Claude Opus 4.8 anthropic/claude-opus-4.8 | $5.00/M | $25.00/M | 1M | $11.00/M | |
| Claude Opus 4.7 anthropic/claude-opus-4.7 | $5.00/M | $25.00/M | 1M | $11.00/M | |
| Claude Opus 5 anthropic/claude-opus-5 | $5.00/M | $25.00/M | 1M | $11.00/M | |
| Claude Fable 5 (batch) anthropic/claude-fable-5:batch | $5.00/M | $25.00/M | 1M | $11.00/M | |
| Claude Opus 4.5 anthropic/claude-opus-4.5 | $5.00/M | $25.00/M | 200K | $11.00/M | |
| Grok 3 Fast Grok 3 Fast | $5.00/M | $25.00/M | 131K | $11.00/M | |
| GPT-5.5 openai/gpt-5.5 | $5.00/M | $30.00/M | 1.1M | $12.50/M | |
| GPT-5.6 Sol Pro openai/gpt-5.6-sol-pro | $5.00/M | $30.00/M | 1.1M | $12.50/M | |
| GPT-5.6 Sol GPT-5.6 Sol | $5.00/M | $30.00/M | 1M | $12.50/M | |
| Fugu Ultra sakana/fugu-ultra | $5.00/M | $30.00/M | 1M | $12.50/M | |
| GPT Chat Latest openai/gpt-chat-latest | $5.00/M | $30.00/M | 400K | $12.50/M | |
| o1 (batch) openai/o1:batch | $7.50/M | $30.00/M | 200K | $14.25/M | |
| GPT-4 Turbo openai/gpt-4-turbo | $10.00/M | $30.00/M | 128K | $16.00/M | |
| GPT-4 Turbo Preview openai/gpt-4-turbo-preview | $10.00/M | $30.00/M | 128K | $16.00/M | |
| Claude Opus 4.1 (batch) anthropic/claude-opus-4.1:batch | $7.50/M | $37.50/M | 200K | $16.50/M | |
| o3 Deep Research openai/o3-deep-research | $10.00/M | $40.00/M | 200K | $19.00/M | |
| o3 Pro (batch) openai/o3-pro:batch | $10.00/M | $40.00/M | 200K | $19.00/M | |
| Claude Fable 5 anthropic/claude-fable-5 | $10.00/M | $50.00/M | 1M | $22.00/M | |
| Claude Opus 5 (Fast) anthropic/claude-opus-5-fast | $10.00/M | $50.00/M | 1M | $22.00/M | |
| Claude Opus 4.8 (Fast) anthropic/claude-opus-4.8-fast | $10.00/M | $50.00/M | 1M | $22.00/M | |
| GPT-5 Pro (batch) openai/gpt-5-pro:batch | $7.50/M | $60.00/M | 400K | $23.25/M | |
| o1 openai/o1 | $15.00/M | $60.00/M | 200K | $28.50/M | |
| GPT-5.2 Pro (batch) openai/gpt-5.2-pro:batch | $10.50/M | $84.00/M | 400K | $32.55/M | |
| Claude Opus 4.1 anthropic/claude-opus-4.1 | $15.00/M | $75.00/M | 200K | $33.00/M | |
| Claude Opus 4 anthropic/claude-opus-4 | $15.00/M | $75.00/M | 200K | $33.00/M | |
| GPT-5.5 Pro (batch) openai/gpt-5.5-pro:batch | $15.00/M | $90.00/M | 1.1M | $37.50/M | |
| GPT-5.4 Pro (batch) openai/gpt-5.4-pro:batch | $15.00/M | $90.00/M | 1.1M | $37.50/M | |
| o3 Pro openai/o3-pro | $20.00/M | $80.00/M | 200K | $38.00/M | |
| GPT-4 openai/gpt-4 | $30.00/M | $60.00/M | 8K | $39.00/M | |
| GPT-5 Pro openai/gpt-5-pro | $15.00/M | $120.00/M | 400K | $46.50/M | |
| GPT-5.2 Pro openai/gpt-5.2-pro | $21.00/M | $168.00/M | 400K | $65.10/M | |
| Claude Opus 4.7 (Fast) anthropic/claude-opus-4.7-fast | $30.00/M | $150.00/M | 1M | $66.00/M | |
| GPT-5.4 Pro openai/gpt-5.4-pro | $30.00/M | $180.00/M | 1.1M | $75.00/M | |
| GPT-5.5 Pro openai/gpt-5.5-pro | $30.00/M | $180.00/M | 1.1M | $75.00/M | |
| o1-pro (batch) openai/o1-pro:batch | $75.00/M | $300.00/M | 200K | $142.50/M | |
| o1-pro openai/o1-pro | $150.00/M | $600.00/M | 200K | $285.00/M |
AI model pricing is quoted per million tokens, split into input (the prompt) and output (the completion). Across the models Tokenando tracks, input prices range from free to over $100 per million tokens, with most frontier models between $1 and $15 per million. Tokenando lists input, output, and a blended 70/30 reference rate for every model.
A blended price is a single comparison number combining input and output rates. Tokenando's default blended rate weights input at 70% and output at 30%, (input × 0.7) + (output × 0.3), reflecting a typical short-prompt, longer-completion workload. Always check input and output separately for your own traffic mix.
The cheapest option depends on your workload. Several providers offer free-tier models, and among paid models the lowest blended rates often come from efficient small models and open-weights models hosted on inference platforms. Sort the index by blended cost to see the current cheapest models.
Pricing is updated on a rolling basis as providers publish changes, and this page refreshes from the latest snapshot every few minutes. Prices are aggregated from live marketplace data and hand-verified against official provider pricing pages.