Llama-3.1-Nemotron-Ultra-253B
Ultra
Nemotron Ultra, NVIDIA's flagship Nemotron tuned over Llama 3.1. 253B parameter dense model targeting frontier-class quality.
Llama-3.1-Nemotron-Ultra-253B is a frontier AI model from Nvidia. It costs $1.600 per million input tokens and $1.600 per million output tokens (blended $1.600/M), with a 128,000-token context window.
INPUT
$1.600/M
per million input tokens
OUTPUT
$1.600/M
per million output tokens
CONTEXT
128,000
tokens
What it is good at
- Largest Nemotron
- Open weights
- NVIDIA NIM hosting
Typical use cases
- Frontier-class self-hosted inference
- Synthetic data generation
Benchmarks
vs. best public score
Hand-curated from each provider's published reports and public leaderboards. Methodology varies across sources, treat as directional rather than authoritative.
More from Nvidia
See all 6 →Llama-3.1-Nemotron-70B
RLHF-Aligned · 128,000 ctx
in $0.350/Mout $0.400/M
Nemotron-4-340B
Large · 4,000 ctx
in $4.200/Mout $4.200/M
Mistral-NeMo-12B (NIM)
Efficient · 128,000 ctx
in $0.150/Mout $0.150/M
Phi-3-Mini-4K (NIM)
Nano · 4,000 ctx
in $0.040/Mout $0.040/M
Mistral-Large-2 (NIM)
Enterprise · 128,000 ctx
in $2.000/Mout $6.000/M
Frequently asked questions
How much does Llama-3.1-Nemotron-Ultra-253B cost?
Llama-3.1-Nemotron-Ultra-253B costs $1.600 per million input tokens and $1.600 per million output tokens, for a blended reference rate of $1.600 per million tokens.
What is Llama-3.1-Nemotron-Ultra-253B's context window?
Llama-3.1-Nemotron-Ultra-253B supports up to 128,000 tokens of context in a single request.
What is Llama-3.1-Nemotron-Ultra-253B best for?
Llama-3.1-Nemotron-Ultra-253B is well suited to Largest Nemotron, Open weights and NVIDIA NIM hosting.
Who makes Llama-3.1-Nemotron-Ultra-253B?
Llama-3.1-Nemotron-Ultra-253B is developed and served by Nvidia.