DeepSeek-V4-Flash
The public-beta release of DeepSeek V4 Flash, which supersedes the earlier preview with substantially stronger agentic behaviour; DeepSeek reports its agent scores beating the larger V4-Pro-Preview. Its defining feature is the price/capability ratio: a 1M-token context window at a fraction of frontier rates. The listed price is DeepSeek's cache-miss rate; cache hits bill lower.
DeepSeek-V4-Flash is a efficient AI model from DeepSeek. It costs $0.140 per million input tokens and $0.280 per million output tokens (blended $0.182/M), with a 1,000,000-token context window.
- Very low cost for a 1M-context model
- Agentic task performance
- Long-context handling
- Cheaper cache-hit pricing
- Open-weight lineage
- High-volume agent loops
- Long-document processing at scale
- Cost-constrained production workloads
- Bulk analysis pipelines
Benchmarks
More from DeepSeek
See all 11 →Frequently asked questions
How much does DeepSeek-V4-Flash cost?
DeepSeek-V4-Flash costs $0.140 per million input tokens and $0.280 per million output tokens, for a blended reference rate of $0.182 per million tokens.
What is DeepSeek-V4-Flash's context window?
DeepSeek-V4-Flash supports up to 1,000,000 tokens of context in a single request.
What is DeepSeek-V4-Flash best for?
DeepSeek-V4-Flash is well suited to Very low cost for a 1M-context model, Agentic task performance and Long-context handling.
Who makes DeepSeek-V4-Flash?
DeepSeek-V4-Flash is developed and served by DeepSeek. It was released in Jul 2026.