Claude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/MClaude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/M
Lepton AI
Lepton AI
Efficient

Llama 3.1 70B (Lepton)

Serverless

The 2024 70B Llama that defined the open-weights chat baseline before Llama 3.3. Still common where hosts haven't upgraded yet.

Llama 3.1 70B (Lepton) is a efficient AI model from Lepton AI. It costs $0.800 per million input tokens and $0.800 per million output tokens (blended $0.800/M), with a 128,000-token context window.

Profile inherited from upstream Llama 3.1 70B — this is a hosted variant of the same open-weights model.

INPUT
$0.800/M
per million input tokens
OUTPUT
$0.800/M
per million output tokens
BLENDED 70/30
$0.800/M
unchanged since 3 May
CONTEXT
128,000
tokens
What it is good at
  • Open weights
  • 128K context
  • Wide hosted availability
Typical use cases
  • Self-hosted chat
  • Fine-tune base
  • Cost benchmarking

Benchmarks

vs. best public score
Scores inherited from Llama 3.1 70B — this is a hosted variant of the same open-weights model, so the underlying benchmark scores are identical.
MMLU83%
Multitask academic knowledge across 57 subjects.
Graduate-level science questions, "Google-proof".
MATH68%
High-school competition math problems.
Python function synthesis from docstrings.
LMArena Elo1248 Elo
Crowd-sourced head-to-head preference Elo rating.
Hand-curated from each provider's published reports and public leaderboards. Methodology varies across sources, treat as directional rather than authoritative.

More from Lepton AI

See all 3

Frequently asked questions

How much does Llama 3.1 70B (Lepton) cost?

Llama 3.1 70B (Lepton) costs $0.800 per million input tokens and $0.800 per million output tokens, for a blended reference rate of $0.800 per million tokens.

What is Llama 3.1 70B (Lepton)'s context window?

Llama 3.1 70B (Lepton) supports up to 128,000 tokens of context in a single request.

What is Llama 3.1 70B (Lepton) best for?

Llama 3.1 70B (Lepton) is well suited to Open weights, 128K context and Wide hosted availability.

Who makes Llama 3.1 70B (Lepton)?

Llama 3.1 70B (Lepton) is developed and served by Lepton AI. It was released in Jul 2024.

Terms used on this page