Llama-3.1-Nemotron-70B
NVIDIA's instruction-tuned Llama 3.1 70B. RLHF'd on top of Meta's base; consistently outranks vanilla 3.1 70B on chat benchmarks.
Llama-3.1-Nemotron-70B is a frontier AI model from Nvidia. It costs $0.350 per million input tokens and $0.400 per million output tokens (blended $0.365/M), with a 128,000-token context window.
- Improved over base Llama 3.1 70B
- NVIDIA NIM endpoints
- Open weights
- Self-hosted chat (Llama-class)
- NIM-based deployments
- Fine-tune base
Benchmarks
More from Nvidia
See all 6 →Frequently asked questions
How much does Llama-3.1-Nemotron-70B cost?
Llama-3.1-Nemotron-70B costs $0.350 per million input tokens and $0.400 per million output tokens, for a blended reference rate of $0.365 per million tokens.
What is Llama-3.1-Nemotron-70B's context window?
Llama-3.1-Nemotron-70B supports up to 128,000 tokens of context in a single request.
What is Llama-3.1-Nemotron-70B best for?
Llama-3.1-Nemotron-70B is well suited to Improved over base Llama 3.1 70B, NVIDIA NIM endpoints and Open weights.
Who makes Llama-3.1-Nemotron-70B?
Llama-3.1-Nemotron-70B is developed and served by Nvidia. It was released in Oct 2024.