Llama 3.1 8B (Together)
Small, fast Llama 3.1. The default cheap open-weights model in 2024–25 across Groq, Cerebras and Together.
Llama 3.1 8B (Together) is a efficient AI model from Together AI. It costs $0.100 per million input tokens and $0.100 per million output tokens (blended $0.100/M), with a 128,000-token context window.
Profile inherited from upstream Llama 3.1 8B ↗ — this is a hosted variant of the same open-weights model.
- Cheap
- Fast on LPU/Cerebras
- Open weights
- 128K context
- Bulk classification
- Cheap chat
- Edge inference
Benchmarks
More from Together AI
See all 10 →Frequently asked questions
How much does Llama 3.1 8B (Together) cost?
Llama 3.1 8B (Together) costs $0.100 per million input tokens and $0.100 per million output tokens, for a blended reference rate of $0.100 per million tokens.
What is Llama 3.1 8B (Together)'s context window?
Llama 3.1 8B (Together) supports up to 128,000 tokens of context in a single request.
What is Llama 3.1 8B (Together) best for?
Llama 3.1 8B (Together) is well suited to Cheap, Fast on LPU/Cerebras and Open weights.
Who makes Llama 3.1 8B (Together)?
Llama 3.1 8B (Together) is developed and served by Together AI. It was released in Jul 2024.