Claude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/MClaude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/M
DeepSeek
DeepSeek
ReasoningLIVE INDEX

R1 Distill Llama 70B

text->text

Open-weights reasoning model with chain-of-thought trained via RL. Comparable to o1 on math benchmarks at a fraction of the price.

R1 Distill Llama 70B is a reasoning AI model from DeepSeek. It costs $0.800 per million input tokens and $0.800 per million output tokens (blended $0.800/M), with a 8,192-token context window.

Profile inherited from upstream DeepSeek-R1 — this is a hosted variant of the same open-weights model.

INPUT
$0.800/M
per million input tokens
OUTPUT
$0.800/M
per million output tokens
BLENDED 70/30
$0.800/M

0.0%over 76 days · hover to read

R1 Distill Llama 70B — blended price

Reconstructed from 76 days of recorded rate-card changes. Prices hold flat between changes because that is what a posted price does — no value here is interpolated.

CONTEXT
8,192
tokens
What it is good at
  • Open-weights reasoning
  • Visible chain-of-thought
  • Math & code
  • Cheap
Typical use cases
  • Reasoning agents
  • Math/STEM tools
  • Self-hosted reasoning

Benchmarks

vs. best public score
Scores inherited from DeepSeek-R1 — this is a hosted variant of the same open-weights model, so the underlying benchmark scores are identical.
MMLU90%
Multitask academic knowledge across 57 subjects.
Graduate-level science questions, "Google-proof".
MATH95%
High-school competition math problems.
Python function synthesis from docstrings.
Real GitHub issues solved end-to-end.
LMArena Elo1361 Elo
Crowd-sourced head-to-head preference Elo rating.
Hand-curated from each provider's published reports and public leaderboards. Methodology varies across sources, treat as directional rather than authoritative.

More from DeepSeek

See all 11

Frequently asked questions

How much does R1 Distill Llama 70B cost?

R1 Distill Llama 70B costs $0.800 per million input tokens and $0.800 per million output tokens, for a blended reference rate of $0.800 per million tokens.

What is R1 Distill Llama 70B's context window?

R1 Distill Llama 70B supports up to 8,192 tokens of context in a single request.

What is R1 Distill Llama 70B best for?

R1 Distill Llama 70B is well suited to Open-weights reasoning, Visible chain-of-thought and Math & code.

Who makes R1 Distill Llama 70B?

R1 Distill Llama 70B is developed and served by DeepSeek.

Terms used on this page