Claude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/MClaude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/M
DeepSeek
DeepSeek
Reasoning

DeepSeek-R1-Distill-70B

Distilled

R1 reasoning capability distilled into a Llama 3.1 70B base. The "open-weights o1-mini", strong STEM at single-machine deployable size.

DeepSeek-R1-Distill-70B is a reasoning AI model from DeepSeek. It costs $0.350 per million input tokens and $0.880 per million output tokens (blended $0.509/M), with a 128,000-token context window.

INPUT
$0.350/M
per million input tokens
OUTPUT
$0.880/M
per million output tokens
BLENDED 70/30
$0.509/M
unchanged since 3 May
CONTEXT
128,000
tokens
What it is good at
  • Open-weights reasoning at 70B
  • Visible chain-of-thought
  • Strong math/code
Typical use cases
  • Self-hosted reasoning agents
  • STEM tutoring
  • Reasoning research

Benchmarks

vs. best public score
MMLU84%
Multitask academic knowledge across 57 subjects.
Graduate-level science questions, "Google-proof".
MATH92%
High-school competition math problems.
Python function synthesis from docstrings.
Hand-curated from each provider's published reports and public leaderboards. Methodology varies across sources, treat as directional rather than authoritative.

More from DeepSeek

See all 11

Frequently asked questions

How much does DeepSeek-R1-Distill-70B cost?

DeepSeek-R1-Distill-70B costs $0.350 per million input tokens and $0.880 per million output tokens, for a blended reference rate of $0.509 per million tokens.

What is DeepSeek-R1-Distill-70B's context window?

DeepSeek-R1-Distill-70B supports up to 128,000 tokens of context in a single request.

What is DeepSeek-R1-Distill-70B best for?

DeepSeek-R1-Distill-70B is well suited to Open-weights reasoning at 70B, Visible chain-of-thought and Strong math/code.

Who makes DeepSeek-R1-Distill-70B?

DeepSeek-R1-Distill-70B is developed and served by DeepSeek. It was released in Jan 2025.

Terms used on this page