DeepSeek: DeepSeek V4.1 Flash API
DeepSeek-V4.1-Flash is a multimodal Mixture-of-Experts model with 552B backbone parameters that natively processes images and text at up to 1M token context. Its Causal Encoder-Decoder architecture activates only 8B parameters during prefill and 16B during decode, cutting the KV cache footprint to roughly a quarter of DeepSeek-V4-Flash for cost-efficient agentic workloads.
- Context window: 1,040,000 tokens
- Max output: 384,000 tokens
- Input: text
- Output: text
- Reasoning: Supported
- Tool calling: Supported
- File input: Supported
- Released: 2026-09-09
Frequently Asked Questions
What is the context window of DeepSeek V4.1 Flash?
DeepSeek V4.1 Flash supports a context window of up to 1,040,000 tokens.
Does DeepSeek V4.1 Flash support function calling?
Yes. DeepSeek V4.1 Flash supports tool / function calling.
Does DeepSeek V4.1 Flash support reasoning?
Yes. DeepSeek V4.1 Flash is a reasoning-capable model.