Nemotron 3.5 Lightning
text->text
Nemotron 3.5 Lightning is a general-purpose chat-tuned model. Available via Nvidia. 262,144-token context. Budget pricing.
Nemotron 3.5 Lightning is a efficient AI model from Nvidia. It costs $0.100 per million input tokens and $0.250 per million output tokens (blended $0.145/M), with a 262,144-token context window.
INPUT
$0.100/M
per million input tokens
OUTPUT
$0.250/M
per million output tokens
CONTEXT
262,144
tokens
What it is good at
- Solid general chat performance
- Reasonable instruction following
Typical use cases
- General chat
- Drafting and rewriting
- Personal productivity
Benchmarks
No published benchmark scores tracked for this model yet. Frontier and reasoning models from major providers have scores; smaller models and inference-host variants typically inherit the underlying open-weights score.
More from Nvidia
See all 6 →Llama-3.1-Nemotron-70B
RLHF-Aligned · 128,000 ctx
in $0.350/Mout $0.400/M
Llama-3.1-Nemotron-Ultra-253B
Ultra · 128,000 ctx
in $1.600/Mout $1.600/M
Nemotron-4-340B
Large · 4,000 ctx
in $4.200/Mout $4.200/M
Mistral-NeMo-12B (NIM)
Efficient · 128,000 ctx
in $0.150/Mout $0.150/M
Phi-3-Mini-4K (NIM)
Nano · 4,000 ctx
in $0.040/Mout $0.040/M
Mistral-Large-2 (NIM)
Enterprise · 128,000 ctx
in $2.000/Mout $6.000/M
Frequently asked questions
How much does Nemotron 3.5 Lightning cost?
Nemotron 3.5 Lightning costs $0.100 per million input tokens and $0.250 per million output tokens, for a blended reference rate of $0.145 per million tokens.
What is Nemotron 3.5 Lightning's context window?
Nemotron 3.5 Lightning supports up to 262,144 tokens of context in a single request.
What is Nemotron 3.5 Lightning best for?
Nemotron 3.5 Lightning is well suited to Solid general chat performance, Reasonable instruction following and General chat.
Who makes Nemotron 3.5 Lightning?
Nemotron 3.5 Lightning is developed and served by Nvidia.