DeepSeek: DeepSeek Flash Latest
About this model
DeepSeek Flash Latest (DeepSeek-V4.1-Flash) is DeepSeek's current default Flash-tier model on the API under the name deepseek-flash, superseding DeepSeek-V4-Flash and retired vision experiment routes. It is a native multimodal mixture-of-experts model with 552 billion backbone parameters, a one-million-token context window, and MIT-licensed open weights on Hugging Face.
The architecture combines a Causal Encoder-Decoder (CED) stack with Compressed Sparse Attention 2 (CSA2), aggressive KV-cache compression (about 890 bytes per token), Engram conditional memory, and DSpark speculative decoding. It activates roughly 8B parameters per token during prefill and 16B during decode, targeting input-heavy agent workloads with lower memory and bandwidth than prior V4-Flash generations while adding built-in image understanding via DeepSeek-ViT.
Post-training emphasizes controllable reasoning effort (1-100) and large-scale agentic RL. At maximum reasoning effort, official evaluations report strong agentic coding results (including Terminal-Bench 2.1 and DeepSWE v1.1) and competitive reasoning scores such as GPQA Diamond and an internal Codeforces rating, with downloadable weights and API access for production multimodal and coding-agent use.
Benchmark Scores
Technical Specs
- Parameters: 552.0B
- Architecture: MoE Transformer (CED)
- Context Window: 1,000,000 tokens
- Input Modalities: text, image
Hardware Requirements
- VRAM: 510.0 GB
- Compute: Multi-GPU server (~512GB VRAM for FP8 weights)
Pricing
| Input | Output | Currency |
|---|---|---|
| 0.02 / 1M tokens | 1.20 / 1M tokens | USD |