Gemma 3 9B
Open
Mid-size Gemma 3, fits on a single consumer GPU while keeping vision support.
Gemma 3 9B is a efficient AI model from Google. It costs $0.090 per million input tokens and $0.090 per million output tokens (blended $0.090/M), with a 128,000-token context window.
INPUT
$0.090/M
per million input tokens
OUTPUT
$0.090/M
per million output tokens
CONTEXT
128,000
tokens
What it is good at
- Single-GPU deployable
- Vision
- Open weights
Typical use cases
- On-device / on-prem chat
- Fine-tune base
Benchmarks
vs. best public score
Hand-curated from each provider's published reports and public leaderboards. Methodology varies across sources, treat as directional rather than authoritative.
More from Google
See all 24 →Gemini 3.5 Pro
Frontier · 1,000,000 ctx
in $2.000/Mout $12.000/M
Gemini 3.6 Flash
Balanced · 1,050,000 ctx
in $1.500/Mout $7.500/M
Gemini 3.5 Flash
Balanced · 1,000,000 ctx
in $1.500/Mout $9.000/M
Gemini 3.1 Pro
Frontier · 1,000,000 ctx
in $2.000/Mout $12.000/M
Gemini 3 Pro
Frontier · 1,000,000 ctx
in $2.000/Mout $12.000/M
Gemini 3 Flash
Efficient · 1,000,000 ctx
in $0.500/Mout $3.000/M
Frequently asked questions
How much does Gemma 3 9B cost?
Gemma 3 9B costs $0.090 per million input tokens and $0.090 per million output tokens, for a blended reference rate of $0.090 per million tokens.
What is Gemma 3 9B's context window?
Gemma 3 9B supports up to 128,000 tokens of context in a single request.
What is Gemma 3 9B best for?
Gemma 3 9B is well suited to Single-GPU deployable, Vision and Open weights.
Who makes Gemma 3 9B?
Gemma 3 9B is developed and served by Google.