Index methodology

Tokenando Indices

Version 1.6.0. Everything below (formulas, constants, source weights, exclusion rules) renders from the same versioned configuration the settlement job computes with. If a value changes, the version changes, the changelog records it, and already-settled history stays frozen.

Principles

Raw-first: every collector stores the untouched upstream payload before any parsing; settlement reads only stored raw data, never live sources.

Deterministic and idempotent: re-running a settled day must reproduce every value exactly, or the run fails loudly. Settled values are immutable at the database level.

Drop, do not fudge: a lab or GPU with insufficient coverage is omitted for the day with a recorded reason. Nothing is interpolated, carried forward, or estimated.

Everything disclosed: offered prices are labelled offered; volume proxies are named with their bias; every methodology change is versioned and logged.

TTPI: Token Price Index

TTPI(lab, day)      = Σ spend(model) ÷ Σ billed_tokens(model)   over the lab’s covered models   (headline, realized)
blended(model)      = input_price × 0.75 + output_price × 0.25
TTPI_list(lab, day) = Σ blended(model) × volume_share(model)   over the lab’s covered models

The headline is the realized price: the US dollars buyers paid for a lab’s tokens on the day, divided by the tokens they bought, per million. It comes from the spend and token counts OpenRouter reports for each model, so caching discounts and the real mix of input and output tokens are in the number because they are in the bill. Tokens bought with the buyer’s own provider key (BYOK) are excluded, because OpenRouter does not bill them at the model price. The realized series starts on 2026-09-21, the first day that spend was reported; it is not estimated for earlier days.

The list series prices the same models at their posted rates instead, blending input at 75% and output at 25% with no caching, and weights them by volume. It runs from index inception and was TTPI’s only series until methodology 1.5.0. Routed traffic is far more input-heavy than that blend and mostly cached, so the list series runs well above the realized price; the gap between them is what caching and the traffic mix take off the rate card.

Prices come from the OpenRouter models API, captured at 23:45 UTC as the day’s close; prices are converted to USD per million tokens, and per-model price overrides shadow base prices where present. Free variants and zero-priced listings are excluded: TTPI is a paid-token index.

Volume weights come from the OpenRouter rankings snapshot captured the following morning (05:10/05:40 UTC), restricted to rows upstream attributes to the settled day, variants standard and thinking only (free and batch are excluded). Both series use the same rows, models and floors; a model counts as covered for the realized series only when spend was reported for it. This weight source reflects routed traffic, not any lab’s total global volume: direct enterprise API traffic is invisible to every public index. That bias is the honest cost of an open methodology, and it is disclosed here rather than hidden.

A lab publishes only when it has at least 2 models carrying both a price and volume, and when the priced models cover at least 90% of the lab’s included-variant volume. Below either floor, the lab drops for the day with a recorded reason.

LabSlugOpenRouter prefixesClosed-weight
Anthropicanthropicanthropicyes
OpenAIopenaiopenaiyes
Googlegooglegoogleyes
xAIxaix-aiyes
DeepSeekdeepseekdeepseekno
Metametameta-llamano
Mistralmistralmistralaino
Qwenqwenqwenno
Moonshot AImoonshotmoonshotaino
Z.ai (Zhipu)zhipuz-aino
MiniMaxminimaxminimaxno
Tencenttencenttencentno

TPPI: Tier Price Index

member(model)        = matches a line in the published tier registry (first match wins)
TPPI(tier, day)      = Σ spend(model) ÷ Σ billed_tokens(model)   over the tier’s members   (headline, realized)
TPPI_list(tier, day) = Σ blended(model) × volume_share(model)    over the tier’s members   (closed-weight tiers only)

TPPI is the Tier Price Index: one price per class of model, across labs. The classes are Frontier, Production, Open flagship, Everyday, Fast & Quick and Value, as defined for the Follow the Token workshop on 21 September 2026. TTPI follows each lab across all of its models; TPPI follows each class across labs, so it answers what a frontier or an everyday token costs, wherever it is bought.

Membership is a published list of product lines, each placed in one tier, never a price band. A price band would be circular: a model that raised its price would leave its tier, so a tier could only ever look cheap. A line covers every generation of one product (Claude Opus 4.6, 5 and 5.5), so when traffic moves to a cheaper new generation the tier’s price falls, which is how labs mostly cut prices. The workshop named Claude Fable and GPT Astra (Frontier), Claude Opus and GPT Sol (Production), Kimi (Open flagship), Claude Sonnet and GPT Terra (Everyday), Claude Haiku (Fast & Quick), and DeepSeek Flash and GPT Luna (Value); the other labs’ lines are placed by how each lab positions them. Placements are reviewed monthly and every change is a new methodology version. Not yet placed: Tencent’s Hunyuan models (HY3, HY4), whose positioning and weights licence are not yet clear, and small open models such as Gemma, gpt-oss, Llama and Mistral Nemo; embedding, image, audio and code-only variants are outside the tiers by design.

The headline is the realized price, TTPI’s formula applied to the tier: the dollars buyers paid for the tier’s tokens on the day divided by the tokens they bought, per million, so caching discounts and the real input/output mix are included. It starts on 2026-09-21, the first day per-model spend was reported. The list series prices the same members at posted rates with TTPI’s 75/25 blend, from index inception, and is published only for tiers whose lines are all closed-weight. An open-weight model’s posted price on OpenRouter is whichever host it routes to that day: between 21 September and 6 October 2026 that moved Kimi’s list price 13% and DeepSeek Flash’s 40% a day, while their realized prices moved 1 to 1.5%.

A tier publishes a series only when at least 2 members from at least 2 labs are covered and they carry at least 90% of the tier’s member volume; otherwise that series drops for the day with a recorded reason. Neither series adjusts for quality: a better generation at the same price shows as no change.

Until methodology 1.6.0, TPPI was the Production Price Index: a chained price index and an average price paid over one curated production tier. Those two series stopped on 2026-10-06; their settled values stay in the database as frozen history.

TierLineLabOpenRouter matchWeights
FrontierClaude FableAnthropic^anthropic\/claude-.*fableclosed
FrontierGPT AstraOpenAI^openai\/gpt-[\d.]+-astraclosed
ProductionClaude OpusAnthropic^anthropic\/claude-.*opusclosed
ProductionGPT SolOpenAI^openai\/gpt-[\d.]+-solclosed
ProductionGemini ProGoogle^google\/gemini-[\d.]+-pro (not image|tts|customtools)closed
ProductionGrokxAI^x-ai\/grok-\d (not fast|mini|code|image|imagine|vision)closed
ProductionQwen MaxQwen^qwen\/qwen[\d.]*-maxclosed
Open flagshipKimiMoonshot AI^moonshotai\/kimi-k\dopen
Open flagshipDeepSeek flagshipDeepSeek^deepseek\/(deepseek-v[\d.]+-pro|deepseek-v3|deepseek-chat-v3|deepseek-r1)open
Open flagshipGLMZ.ai (Zhipu)^z-ai\/glm-[\d.]+ (not flash|air|-\d*v\b|vision)open
Open flagshipMiniMax MMiniMax^minimax\/minimax-m\dopen
EverydayClaude SonnetAnthropic^anthropic\/claude-.*sonnetclosed
EverydayGPT TerraOpenAI^openai\/gpt-[\d.]+-terraclosed
EverydayGemini FlashGoogle^google\/gemini-[\d.]+-flash (not flash-lite|image|tts|live|transcribe)closed
Fast & QuickClaude HaikuAnthropic^anthropic\/claude-.*haikuclosed
Fast & QuickGPT miniOpenAI^openai\/gpt-[\d.o]+-mini (not image|audio|codex|realtime|search|transcribe|tts)closed
ValueDeepSeek FlashDeepSeek^deepseek\/deepseek-v[\d.]+-flash (not vision)open
ValueGPT LunaOpenAI^openai\/gpt-[\d.]+-lunaclosed
ValueGPT nanoOpenAI^openai\/gpt-[\d.o]+-nano (not image|audio|codex|realtime|search|transcribe|tts)closed
ValueGemini Flash-LiteGoogle^google\/gemini-[\d.]+-flash-lite (not image|tts)closed
ValueGLM FlashZ.ai (Zhipu)^z-ai\/glm-[\d.]+-flashopen
ValueQwen FlashQwen^qwen\/qwen[\d.]*-flashopen

TCPI: Compute Price Index

provider_price(gpu) = median of that provider’s per-GPU offers across the day’s captures
TCPI(gpu, day)      = weighted median of provider_price over all providers   (cleared 2.0 · posted 1.0)

GPU offers are captured four times daily (05:30, 11:30, 17:30, 23:30 UTC) from three disclosed sources: the Vast.ai public ask book (one query per GPU, price-ascending so any truncation at the venue’s 64-offer cap is stable and bites the expensive tail), RunPod’s public GPU pricing (secure and community cloud as two distinct providers), and the Shadeform catalogue, which carries posted rates for 19 underlying clouds, each counted as its own provider, only regions with live availability.

All prices are normalised to USD per single GPU-hour (node prices divided by GPU count; Shadeform’s cent-denominated prices divided by 100, cross-verified against posted dollar rates). Offers outside a per-GPU plausibility band are discarded and counted. Every current source posts asked/on-demand rates, so v1 is honestly an OFFERED-price index; the cleared-source weight is declared now so adding a transaction-based source later is a source addition, not a formula change.

A GPU publishes only with offers from at least 3 distinct providers. Interruptible/spot prices are excluded from the headline and published as a separate companion series. The per-provider median-of-offers followed by a cross-provider weighted median means neither a flood of listings from one venue nor a single outlier provider can move the index.

GPU keyPlausibility band $/GPU-hrVast namesRunPod namesShadeform type
H100 SXM0.5 – 15H100 SXMH100 SXMH100 (sxm)
H100 PCIe0.5 – 12H100 PCIEH100 PCIeH100 (pcie)
H2000.8 – 20H200H200 SXMH200
B2001 – 30B200B200B200
A100 SXM40.2 – 10A100 SXM4A100 SXMA100_80G (sxm)
RTX 40900.05 – 5RTX 4090RTX 4090RTX4090
RTX 50900.08 – 8RTX 5090RTX 5090RTX5090

Day attribution & settlement

All days are UTC. For settled day D: GPU offers are captured four times during D; token prices are captured at 23:45 UTC on D; volume weights are captured at ~05:10 UTC on D+1 (the rankings snapshot upstream attributes to D); settlement runs at 06:00 UTC on D+1. Settled values are frozen the moment they are written.

The published record begins at index inception, 2026-08-15. Settlement refuses to produce a value for any earlier day, so the start of the series is a deliberate, enforced boundary rather than wherever data happened to begin. Values computed before inception, while the methodology was still being written, were removed rather than published; the raw observations behind them were kept.

Sources

SourceTypeFeedsNotes
OpenRouter models APIposted pricesTTPI, TPPI pricesofficial public API; per-model USD-per-token
OpenRouter rankingsvenue-reported volumeTTPI, TPPI weights; TTPI realized spendpublic endpoint; spend reported since 2026-09-21; routed-traffic bias disclosed above
Vercel AI Gateway leaderboardsvenue-reported usage sharescollected daily for validation; not yet an index inputCC BY 4.0, attribution to Vercel; token, spend and request shares per lab and model; spend estimated at list prices
Vercel AI Gateway modelsposted pricescollected daily for validation; not yet an index inputofficial list prices incl. cache, fast mode and EU regional prices
OpenRouter per-host endpointsposted prices, per hostcollected daily for validation; not yet an index inputevery host’s price for the 60 highest-volume index models
Hugging Face Inference Providersposted prices, per hostcollected daily for validation; not yet an index inputopen-weight host prices (Together, Fireworks, Baseten, DeepInfra and others)
AWS Bedrock price listofficial regional pricescollected daily for a future regional seriesUS, EU, UAE, Bahrain, Israel; in-region, global, flex, priority and batch tiers
Azure retail pricesofficial regional pricescollected daily for a future regional seriesFoundry Models in US, EU, UAE, Qatar; global, data-zone and regional deployments
LiteLLM, models.devcommunity price databasesquality checks only; never an index inputflag stale or mis-parsed prices before settle
Vast.ai ask bookposted (marketplace asks)TCPIpublic search API; ask book only, no utilisation signal
RunPod pricingpostedTCPIpublic GraphQL; secure + community as two providers; bids → spot series
Shadeform catalogposted (19 clouds)TCPIpublic API; per-region availability gating; cents→USD verified

Free JSON API

The headline series are free, keyless, and CORS-open. Responses carry the methodology version and coverage metadata on every value.

EndpointReturns
/api/v1/catalogIndices, entities, date ranges, version, disclaimer
/api/v1/ttpi?lab=&series=headline|list&from=&to=Token price index, per lab per day ($/MTok); headline is realized, list is posted prices
/api/v1/tppi?tier=&series=headline|list&from=&to=Tier price index, per model class per day ($/MTok); headline is realized, list is posted prices
/api/v1/tcpi?gpu=&series=headline|spotCompute price index, per GPU per day ($/GPU-hr)
/api/v1/labs · /api/v1/gpusRegistries (config-derived)

Changelog

VersionDateChanges
1.6.02026-10-07

TPPI is now the Tier Price Index: one price per class of model, across labs, for six tiers: Frontier, Production, Open flagship, Everyday, Fast & Quick and Value, as defined for the Follow the Token workshop on 21 September 2026. TTPI follows each lab across its models; TPPI now follows each class across labs.

Each tier is a published list of product lines, each line covering every generation of one product. The workshop named Claude Fable and GPT Astra (Frontier), Claude Opus and GPT Sol (Production), Kimi (Open flagship), Claude Sonnet and GPT Terra (Everyday), Claude Haiku (Fast & Quick), and DeepSeek Flash and GPT Luna (Value). Google, xAI, Qwen, DeepSeek, Z.ai and MiniMax lines are placed by how each lab positions them; the full list is on this page.

Like TTPI since 1.5.0, each tier’s headline is the realized price, spend divided by tokens, from 2026-09-21; a list-price series runs from index inception for tiers made entirely of closed-weight models. Open-weight tiers have no list series, because an open model’s posted price on OpenRouter is whichever host it routes to that day: over 21 September to 6 October that moved Kimi 13% and DeepSeek Flash 40% a day, against 1 to 1.5% for their realized prices. A tier needs members from at least two labs, so it is never one lab’s price.

The Production Price Index series (a chained price index and an average price paid over one production tier) stopped on 2026-10-06. Their settled values were not changed or deleted; they remain in the database as frozen history and are no longer shown or served. Tier history was computed from the stored daily snapshots. TTPI, TCPI and TIS are unchanged.

1.5.02026-10-06

TTPI’s headline is now the realized price: what buyers actually paid per million tokens, from the USD spend and token counts OpenRouter reports for each model each day. Until now TTPI priced tokens at posted list rates with a fixed 75% input, 25% output blend and no caching. Real routed traffic is about 97% input tokens, and most of that input is served from cache at a tenth of the input price or less, so the list figure ran two to five times above what was paid. On 2026-10-04 Anthropic’s list figure was $7.60 per million tokens; its realized price was $1.09.

The realized series starts on 2026-09-21, the first day OpenRouter reported spend per model. It is not estimated for earlier days. It covers the same models as before (those with a posted token price), uses the same volume rows, variants and coverage floors, and excludes tokens bought with the buyer’s own provider key (BYOK), which OpenRouter does not bill at the model price.

The list-price series continues unchanged, published beside the headline as the list series from index inception (2026-08-15). Its settled history was relabelled from headline to list in migration 0013; no value, day or version on any row changed. Reading the two together shows how much caching and the input-heavy mix take off the rate card.

TIS and Frontier Premium are computed from the list series, as before, so they are unchanged. TCPI and TPPI are unchanged.

1.4.02026-10-05

TPPI price index restated. It now chains product lines (families) rather than individual models: a line’s price is the usage-weighted price of its generations, so traffic moving to a cheaper new generation of the same line counts as a price cut. Under 1.3.0 a new generation entered at its launch price as no change, so when GPT-6 Luna launched at $0.20 beside GPT-5.6 Luna at $0.45 the index did not move, and over time it could only register price rises on old models. Restated, the index stands at 109.8 on 2026-10-03 where 1.3.0 showed 157.6.

Two families that combined separate products were split, so the index never reads a move between products as a price change: Gemini Flash and Gemini Flash-Lite, GPT mini and GPT nano.

This is a restatement of published values, which the methodology otherwise never does. TPPI had been public for one day under 1.3.0; every TPPI value since inception (2026-08-15) was removed and recomputed from the stored daily snapshots under 1.4.0, recorded in migration 0011. The average price paid is unchanged in value and was re-settled with it, so both series carry one version. TTPI, TCPI, TIS and Frontier Premium are unchanged.

1.3.02026-09-25

New index: TPPI, the Tokenando Production Price Index. It tracks the production tier (Sonnet, Haiku, Luna, Flash, Gemma, Llama, Qwen Flash and similar) across labs, where TTPI tracks each lab across all of its tiers. Frontier flagships have converged on the same posted price; the production tier is where labs compete on price.

Membership is a curated list of named tiers with a published rule, reviewed monthly, never a price band: a band would drop any model that raised its price, so the index could only ever look cheap. Open-weight families are included. Labs that sell a single tier are out.

TPPI publishes two series. The first is a chained price index (base 100) over closed-weight members that moves only when those models change price. Open-weight members are left out of it, because their OpenRouter price is the cheapest host of the day and flips daily: combined, that host churn moved the index ±15% a day, and an open-only index moved 20-50% a day. An open-weight index is withheld until open models are priced across all their hosts. The second series is the volume-weighted average price paid over all members, the TTPI formula, which also moves when traffic shifts between cheaper and dearer models. Reading them together separates price cuts from adoption of cheaper models.

Existing indices are unchanged: TTPI, TCPI, TIS and Frontier Premium compute exactly as under 1.2.2. TPPI history was computed from the stored daily snapshots back to index inception.

1.2.22026-08-16

Correction to the 1.2.1 note below: it said the two construction days had token prices captured "roughly a day late". That was accurate for 2026-08-14 (22h57m late) but overstated for 2026-08-15, whose prices were 7h25m late. Posted token prices change on announcement, roughly monthly, so a seven-hour offset does not move the index materially.

Index inception accordingly moved from 2026-08-17 to 2026-08-15, and that day settled. It is the first day with a complete input set: GPU offers captured within the day itself, and volume weights captured the following morning exactly as designed. The late price capture is disclosed on the snapshot record rather than smoothed over.

2026-08-14 remains excluded permanently: its prices were nearly a full day late and no GPU offers were collected at all, so two of the three indices could never settle for it.

1.2.12026-08-16

Index inception set to 2026-08-17: settlement now refuses any earlier day outright, so the published record has a defined and enforced start.

Values settled during construction on 2026-08-14 and 2026-08-15 were removed (migration 0008). Their token prices were captured roughly a day late, because the end-of-day collector did not exist when those days closed, so the snapshots held the following day’s prices, and the methodology changed three times across them, leaving values the current configuration cannot reproduce. Publishing two days that fail our own reproducibility test, at the exact point a reader looks hardest, was not worth two days of history.

Every raw snapshot from those days was kept, immutable and untouched, including volume observations that can never be re-fetched from upstream. Nothing irreplaceable was discarded, and the settle-run log of what happened during construction was kept as an audit trail.

1.2.02026-08-16

TIS expanded from one open model on one GPU to five models across nine model-and-GPU combinations (DeepSeek V4 Pro and R1, Kimi K2.6, MiniMax M3 and gpt-oss 120B on H100 SXM, H200 and B200), each at three serving speeds, all sourced from InferenceX comparison pages with per-pairing citations. Readers choose the combination rather than inheriting ours.

Frontier Premium is now restricted to frontier-class open models. gpt-oss 120B is excluded from it: at roughly 60M tokens per GPU-hour it would set the cheapest self-hosted cost by being small and fast, turning the headline ratio into a statement about model size rather than hosting economics. It remains fully available in TIS.

Pairings that answer no real question are no longer settled: an open-weight lab’s API against somebody else’s open model. Only like-for-like (the lab that serves this exact model) and substitution (closed labs, which cannot be self-hosted at all) are published.

Reproducibility checks are now scoped to the methodology version that produced a value. A value settled under an earlier version is frozen history; it cannot be re-derived from today’s constants and is never overwritten or reported as a mismatch. Without this, the first constant change would have made every historical day fail to re-settle, contradicting the purpose of versioning.

1.1.12026-08-16

Throughput rows now record which lab publishes the open weights, so the index can distinguish a like-for-like pairing (a lab’s own API vs self-hosting its own model, a clean hosting-margin measurement) from a substitution pairing (a closed lab’s API vs switching to a different, open model, which trades away model quality too). The two answer different questions and are now reported separately instead of in one undifferentiated list.

Documented that the implied self-hosting cost assumes full utilisation of the rented GPU and excludes idle capacity, failed requests and operations overhead. It is a floor, not a forecast.

Labelling and presentation only: no computed value changes, and no settled value was altered.

1.1.02026-08-16

TIS now settles three serving operating points per (model, GPU), namely volume, balanced and latency, instead of one. Published benchmark curves span roughly 11x between the batch and low-latency ends, which moves the Frontier Premium by more than the difference between any two labs; picking one point silently was indefensible, so all three publish and the reader can switch.

The "balanced" point (interactive serving, ~116 tokens/sec per user) is the documented headline: the API default, the citable figure, and the number in the daily summary.

Throughput table reseeded from InferenceX (SemiAnalysis) for DeepSeek V4 Pro 1.6T on B200 at FP4, replacing the unverified DeepSeek V3 / H200 placeholder. InferenceX reports H200 as unmeasured for this model. Operating points are read from an interpolated published curve, disclosed as such.

Weekly staleness check added: verified throughput citations older than 90 days, or whose source URL stops resolving, raise an alert rather than ageing silently.

1.0.02026-08-16

Initial methodology: TTPI (volume-weighted blended $/MTok per lab, 0.75/0.25 input/output share), TCPI (weighted median $/GPU-hr per GPU from disclosed posted sources), TIS and Frontier Premium (verified-throughput rows only).

Sources: OpenRouter models API (prices), OpenRouter rankings snapshot (volume weights; disclosed with bias), Vast.ai ask book, RunPod GraphQL, Shadeform catalogue (19 clouds).

Drop rules: labs below 90% volume coverage or under 2 priced models drop for the day; GPUs under 3 distinct providers drop for the day. Missing data is never interpolated or carried forward.

How prices for individual models are sourced and verified is documented separately on the About page.

Pricing-catalogue methodology →

Tokenando indices are informational reference data, not financial products, price quotes, or investment advice. Offered prices are labeled offered; sources, weights, and exclusion rules are fully disclosed in the methodology. Settled values are immutable; corrections ship as new methodology versions, never as silent edits.