Reka Flash 21B
Efficient
Faster, cheaper Reka Flash. 21B parameters, multimodal, deployable on a single high-end GPU.
Reka Flash 21B is a efficient AI model from Reka AI. It costs $0.800 per million input tokens and $4.000 per million output tokens (blended $1.760/M), with a 128,000-token context window.
INPUT
$0.800/M
per million input tokens
OUTPUT
$4.000/M
per million output tokens
CONTEXT
128,000
tokens
What it is good at
- Single-GPU multimodal
- Native vision + audio
- Cheaper than Reka Core
Typical use cases
- Self-hosted multimodal chat
- Cost-sensitive video QA
Benchmarks
vs. best public score
Hand-curated from each provider's published reports and public leaderboards. Methodology varies across sources, treat as directional rather than authoritative.
More from Reka AI
See all 3 →Frequently asked questions
How much does Reka Flash 21B cost?
Reka Flash 21B costs $0.800 per million input tokens and $4.000 per million output tokens, for a blended reference rate of $1.760 per million tokens.
What is Reka Flash 21B's context window?
Reka Flash 21B supports up to 128,000 tokens of context in a single request.
What is Reka Flash 21B best for?
Reka Flash 21B is well suited to Single-GPU multimodal, Native vision + audio and Cheaper than Reka Core.
Who makes Reka Flash 21B?
Reka Flash 21B is developed and served by Reka AI.