Claude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/MClaude Fable 5$22.000/MClaude Opus 5$11.000/MClaude Opus 4.8$11.000/MClaude Opus 4.7$11.000/MClaude Opus 4.6$11.000/MClaude Opus 4.5$33.000/MClaude Sonnet 3.7$6.600/MClaude Opus 3$33.000/MClaude 2.1$12.800/MClaude 2$12.800/MGPT-5.6 Sol$12.500/MGPT-5.6 Terra$5.000/MGPT-5.5$12.500/MGPT-5.2$5.425/MGPT-5.2-Codex$5.425/MGPT-5$3.875/MGPT-4.5$97.500/MGPT-4 Turbo Preview$16.000/MGPT-4$39.000/MGPT-4-32k$78.000/Mo3$19.000/Mo3-mini$2.090/Mo4-mini$2.090/Mo1$28.500/Mo1-mini$5.700/Mo1-preview$28.500/MGemini 3.5 Pro$5.000/MGemini 3.1 Pro$5.000/MGemini 3 Pro$5.000/MGemini 2.5 Pro$3.875/M
Meta
Meta
EfficientLIVE INDEX

Llama 3.3 70B Instruct

text->text

The most-deployed open-weights chat model in 2025, strong reasoning at 70B with broad inference-provider support.

Llama 3.3 70B Instruct is a efficient AI model from Meta. It costs $0.710 per million input tokens and $0.710 per million output tokens (blended $0.710/M), with a 131,072-token context window.

Profile inherited from upstream Llama 3.3 70B — this is a hosted variant of the same open-weights model.

INPUT
$0.710/M
per million input tokens
OUTPUT
$0.710/M
per million output tokens
BLENDED 70/30
$0.710/M
unchanged since 3 May
CONTEXT
131,072
tokens
What it is good at
  • Strong open chat baseline
  • Cheap on Groq/Cerebras
  • 128K context
  • Wide ecosystem
Typical use cases
  • Self-hosted production chat
  • Cost benchmarking
  • RAG

Benchmarks

vs. best public score
Scores inherited from Llama 3.3 70B — this is a hosted variant of the same open-weights model, so the underlying benchmark scores are identical.
MMLU86%
Multitask academic knowledge across 57 subjects.
Graduate-level science questions, "Google-proof".
MATH77%
High-school competition math problems.
Python function synthesis from docstrings.
LMArena Elo1257 Elo
Crowd-sourced head-to-head preference Elo rating.
Hand-curated from each provider's published reports and public leaderboards. Methodology varies across sources, treat as directional rather than authoritative.

More from Meta

See all 20

Frequently asked questions

How much does Llama 3.3 70B Instruct cost?

Llama 3.3 70B Instruct costs $0.710 per million input tokens and $0.710 per million output tokens, for a blended reference rate of $0.710 per million tokens.

What is Llama 3.3 70B Instruct's context window?

Llama 3.3 70B Instruct supports up to 131,072 tokens of context in a single request.

What is Llama 3.3 70B Instruct best for?

Llama 3.3 70B Instruct is well suited to Strong open chat baseline, Cheap on Groq/Cerebras and 128K context.

Who makes Llama 3.3 70B Instruct?

Llama 3.3 70B Instruct is developed and served by Meta.

Terms used on this page