DeepSeek: DeepSeek V4.1 Flash (batch)
About this model
DeepSeek-V4.1-Flash is a multimodal mixture-of-experts language model designed for long-context, agentic, and vision-language workloads. It uses a 40-layer causal encoder-decoder (CED) Transformer with Compressed Sparse Attention 2 (CSA2), FP4 KV caching, Engram conditional memory, and DSpark speculative decoding, activating about 8B parameters during prefill and 16B during decode from a 552B-parameter backbone. The model supports up to one million tokens of context, native image-and-text input, continuously adjustable reasoning effort (1–100), and strong scores on terminal, software-engineering, and cybersecurity agent benchmarks when evaluated at maximum reasoning effort.
The open weights and reference inference code are released on Hugging Face under the MIT license, while the DeepSeek Batch API offers the same V4.1-Flash instruct capabilities for asynchronous, cost-efficient bulk inference. Post-training follows SFT, reinforcement learning, and on-policy distillation with large-scale synthetic agent data. On standard knowledge and reasoning suites the instruct model reaches competitive GPQA-Diamond, LiveBench, and AIME 2025 results; on agent harnesses it leads on Terminal-Bench 2.1, DeepSWE v1.1, CyberGym, and AutomationBench among compared DeepSeek variants.
DeepSeek-V4.1-Flash targets production agents that read long documents or repositories, run coding and terminal tools, and interpret charts or UI screenshots without sacrificing throughput. KV cache compression to roughly 890 bytes per token materially reduces memory for million-token sessions relative to earlier DeepSeek-Flash generations, making it practical for input-heavy pipelines served via API or self-hosted with multi-GPU clusters.
Benchmark Scores
Technical Specs
- Parameters: 552.0B
- Architecture: Mixture-of-Experts (MoE) Transformer
- Context Window: 1,000,000 tokens
- Input Modalities: text, image
Hardware Requirements
- Compute: API only
Pricing
| Input | Output | Currency |
|---|---|---|
| 0.11 / 1M tokens | 0.34 / 1M tokens | USD |