Llama 3.1 70B
Balanced
The 2024 70B Llama that defined the open-weights chat baseline before Llama 3.3. Still common where hosts haven't upgraded yet.
Llama 3.1 70B is a efficient AI model from Meta. It costs $0.350 per million input tokens and $0.400 per million output tokens (blended $0.365/M), with a 128,000-token context window.
INPUT
$0.350/M
per million input tokens
OUTPUT
$0.400/M
per million output tokens
CONTEXT
128,000
tokens
What it is good at
- Open weights
- 128K context
- Wide hosted availability
Typical use cases
- Self-hosted chat
- Fine-tune base
- Cost benchmarking
Benchmarks
vs. best public score
Hand-curated from each provider's published reports and public leaderboards. Methodology varies across sources, treat as directional rather than authoritative.
More from Meta
See all 20 →Muse Spark 1.2
Frontier · 1,049,000 ctx
in $1.250/Mout $4.250/M
Llama 4 Maverick
Frontier · 128,000 ctx
in $0.200/Mout $0.600/M
Llama 4 Scout
Efficient · 512,000 ctx
in $0.100/Mout $0.350/M
Llama 3.3 70B
Open · 128,000 ctx
in $0.230/Mout $0.400/M
Llama 3.1 405B
Large · 128,000 ctx
in $0.800/Mout $0.800/M
Llama 3.1 8B
Compact · 128,000 ctx
in $0.100/Mout $0.100/M
Frequently asked questions
How much does Llama 3.1 70B cost?
Llama 3.1 70B costs $0.350 per million input tokens and $0.400 per million output tokens, for a blended reference rate of $0.365 per million tokens.
What is Llama 3.1 70B's context window?
Llama 3.1 70B supports up to 128,000 tokens of context in a single request.
What is Llama 3.1 70B best for?
Llama 3.1 70B is well suited to Open weights, 128K context and Wide hosted availability.
Who makes Llama 3.1 70B?
Llama 3.1 70B is developed and served by Meta. It was released in Jul 2024.