Claude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/MClaude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/M
DeepSeek
DeepSeek
Reasoning

DeepSeek-R1-Distill-8B

Distilled

Smallest R1 distill, Llama 3.1 8B base. Tradeoff: limited reasoning depth vs the larger distills.

DeepSeek-R1-Distill-8B is a reasoning AI model from DeepSeek. It costs $0.070 per million input tokens and $0.200 per million output tokens (blended $0.109/M), with a 128,000-token context window.

INPUT
$0.070/M
per million input tokens
OUTPUT
$0.200/M
per million output tokens
BLENDED 70/30
$0.109/M
unchanged since 3 May
CONTEXT
128,000
tokens
What it is good at
  • Smallest R1 distill
  • Edge-friendly
  • Open weights
Typical use cases
  • Edge reasoning experiments
  • Distillation research

Benchmarks

vs. best public score
MMLU67%
Multitask academic knowledge across 57 subjects.
Graduate-level science questions, "Google-proof".
MATH75%
High-school competition math problems.
Python function synthesis from docstrings.
Hand-curated from each provider's published reports and public leaderboards. Methodology varies across sources, treat as directional rather than authoritative.

More from DeepSeek

See all 11

Frequently asked questions

How much does DeepSeek-R1-Distill-8B cost?

DeepSeek-R1-Distill-8B costs $0.070 per million input tokens and $0.200 per million output tokens, for a blended reference rate of $0.109 per million tokens.

What is DeepSeek-R1-Distill-8B's context window?

DeepSeek-R1-Distill-8B supports up to 128,000 tokens of context in a single request.

What is DeepSeek-R1-Distill-8B best for?

DeepSeek-R1-Distill-8B is well suited to Smallest R1 distill, Edge-friendly and Open weights.

Who makes DeepSeek-R1-Distill-8B?

DeepSeek-R1-Distill-8B is developed and served by DeepSeek. It was released in Jan 2025.

Terms used on this page