Llama-3.3-70B (Cerebras)
The most-deployed open-weights chat model in 2025, strong reasoning at 70B with broad inference-provider support.
Llama-3.3-70B (Cerebras) is a efficient AI model from Cerebras. It costs $0.590 per million input tokens and $0.990 per million output tokens (blended $0.710/M), with a 128,000-token context window.
Profile inherited from upstream Llama 3.3 70B ↗ — this is a hosted variant of the same open-weights model.
- Strong open chat baseline
- Cheap on Groq/Cerebras
- 128K context
- Wide ecosystem
- Self-hosted production chat
- Cost benchmarking
- RAG
Benchmarks
More from Cerebras
See all 4 →Frequently asked questions
How much does Llama-3.3-70B (Cerebras) cost?
Llama-3.3-70B (Cerebras) costs $0.590 per million input tokens and $0.990 per million output tokens, for a blended reference rate of $0.710 per million tokens.
What is Llama-3.3-70B (Cerebras)'s context window?
Llama-3.3-70B (Cerebras) supports up to 128,000 tokens of context in a single request.
What is Llama-3.3-70B (Cerebras) best for?
Llama-3.3-70B (Cerebras) is well suited to Strong open chat baseline, Cheap on Groq/Cerebras and 128K context.
Who makes Llama-3.3-70B (Cerebras)?
Llama-3.3-70B (Cerebras) is developed and served by Cerebras.