Gemini 1.5 Flash-8B
Nano
Smallest Gemini 1.5, 8B parameters, distilled for ultra-cheap inference. Punches above its size on routing/classification.
Gemini 1.5 Flash-8B is a efficient AI model from Google. It costs $0.037 per million input tokens and $0.150 per million output tokens (blended $0.071/M), with a 1,000,000-token context window.
INPUT
$0.037/M
per million input tokens
OUTPUT
$0.150/M
per million output tokens
CONTEXT
1,000,000
tokens
What it is good at
- Very cheap
- Sub-100ms latency on most tokens
- 1M context
Typical use cases
- Routing layer
- Bulk classification
- Edge-style inference
Benchmarks
vs. best public score
Hand-curated from each provider's published reports and public leaderboards. Methodology varies across sources, treat as directional rather than authoritative.
More from Google
See all 24 →Gemini 3.5 Pro
Frontier · 1,000,000 ctx
in $2.000/Mout $12.000/M
Gemini 3.6 Flash
Balanced · 1,050,000 ctx
in $1.500/Mout $7.500/M
Gemini 3.5 Flash
Balanced · 1,000,000 ctx
in $1.500/Mout $9.000/M
Gemini 3.1 Pro
Frontier · 1,000,000 ctx
in $2.000/Mout $12.000/M
Gemini 3 Pro
Frontier · 1,000,000 ctx
in $2.000/Mout $12.000/M
Gemini 3 Flash
Efficient · 1,000,000 ctx
in $0.500/Mout $3.000/M
Frequently asked questions
How much does Gemini 1.5 Flash-8B cost?
Gemini 1.5 Flash-8B costs $0.037 per million input tokens and $0.150 per million output tokens, for a blended reference rate of $0.071 per million tokens.
What is Gemini 1.5 Flash-8B's context window?
Gemini 1.5 Flash-8B supports up to 1,000,000 tokens of context in a single request.
What is Gemini 1.5 Flash-8B best for?
Gemini 1.5 Flash-8B is well suited to Very cheap, Sub-100ms latency on most tokens and 1M context.
Who makes Gemini 1.5 Flash-8B?
Gemini 1.5 Flash-8B is developed and served by Google. It was released in Oct 2024.