Claude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/MClaude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/M

Tokenando Indices

TIS — Inference Spread

What a lab charges per million tokens minus what the same tokens cost to serve yourself on rented GPUs with an open-weight model. Throughput assumptions come from cited public benchmarks and publish only after human verification. The Frontier Premium is the headline ratio: cheapest closed-lab price over cheapest credible self-hosted equivalent.

5 open models across 9 model-and-GPU combinations, each at 3 serving speeds · assumptions and citations →

A serving stack can be tuned to push many tokens per GPU (each user sees slower output) or few tokens per GPU (each user sees faster output). Published benchmarks for one model on one GPU span roughly 11× across that curve — more than the gap between two labs — so the assumption you make decides the answer. Rather than pick one quietly, all three settle every day and you can switch between them. "Balanced" is the published default because it matches the responsiveness a commercial API actually delivers.

Interactive serving at roughly the responsiveness a commercial API delivers — the closest like-for-like comparison, and the published headline. DeepSeek V4 Pro 1.6T on a B200 produces 1,687 tokens/sec per GPU at this setting (6.07M per GPU-hour), so serving it yourself costs $1.12 per million tokens — before idle capacity or operations overhead. Frontier Premium at this setting: 0.83×. Source: InferenceX (SemiAnalysis) — DeepSeek V4 Pro 1.6T, B200 vs H200 (FP4, 8k/1k seq).

Hosting margin — the same model, hosted vs self-hosted

Identical weights and identical output quality; the only thing that changes is who runs the GPUs. The spread is purely what you pay for someone else to operate the hardware.

ModelAPI price $/MTokSelf-host $/MTokSpreadCheaperΔ day
DeepSeek vs deepseek-v4-pro on B200$0.209$1.12-$0.909the API+6.83%

Frontier substitution — a closed API vs self-hosting this model

Claude, GPT and Gemini weights are not released at any price, so no like-for-like self-hosted comparison for them can exist. The spread below bundles a price gap with a capability gap — a procurement question, never a like-for-like saving.

LabAPI price $/MTokSelf-host $/MTokSpreadCheaperΔ day
Anthropic vs deepseek-v4-pro on B200$8.17$1.12$7.05self-hosting-0.68%
xAI vs deepseek-v4-pro on B200$2.81$1.12$1.70self-hosting+0.01%
OpenAI vs deepseek-v4-pro on B200$1.92$1.12$0.803self-hosting+3.61%
Google vs deepseek-v4-pro on B200$0.924$1.12-$0.194the API+10.25%

Spread history

No settled history for this combination yet — the first settles fill this chart.

GET /api/v1/tis?model=kimi-k2.6&gpu=H200&profile=balanced · GET /api/v1/frontier-premium?profile=all — free, no key, CORS open.

Tokenando indices are informational reference data, not financial products, price quotes, or investment advice. Offered prices are labeled offered; sources, weights, and exclusion rules are fully disclosed in the methodology. Settled values are immutable; corrections ship as new methodology versions, never as silent edits.