Reka AI
Reka AI
Efficient

Reka Flash 21B

Efficient

Faster, cheaper Reka Flash. 21B parameters, multimodal, deployable on a single high-end GPU.

Reka Flash 21B is a efficient AI model from Reka AI. It costs $0.800 per million input tokens and $4.000 per million output tokens (blended $1.760/M), with a 128,000-token context window.

INPUT
$0.800/M
per million input tokens
OUTPUT
$4.000/M
per million output tokens
BLENDED 70/30
$1.760/M
unchanged since 3 May
CONTEXT
128,000
tokens
What it is good at
  • Single-GPU multimodal
  • Native vision + audio
  • Cheaper than Reka Core
Typical use cases
  • Self-hosted multimodal chat
  • Cost-sensitive video QA

Benchmarks

vs. best public score
MMLU75%
Multitask academic knowledge across 57 subjects.
Python function synthesis from docstrings.
LMArena Elo1184 Elo
Crowd-sourced head-to-head preference Elo rating.
Hand-curated from each provider's published reports and public leaderboards. Methodology varies across sources, treat as directional rather than authoritative.

More from Reka AI

See all 3 →

Frequently asked questions

How much does Reka Flash 21B cost?

Reka Flash 21B costs $0.800 per million input tokens and $4.000 per million output tokens, for a blended reference rate of $1.760 per million tokens.

What is Reka Flash 21B's context window?

Reka Flash 21B supports up to 128,000 tokens of context in a single request.

What is Reka Flash 21B best for?

Reka Flash 21B is well suited to Single-GPU multimodal, Native vision + audio and Cheaper than Reka Core.

Who makes Reka Flash 21B?

Reka Flash 21B is developed and served by Reka AI.

Terms used on this page