AI Agent Hub
Back to models
DeepSeek: DeepSeek V4.1 Flash (batch) logo

DeepSeek: DeepSeek V4.1 Flash (batch)

Multimodal deepseek Released 2025-10-09
39.0 / 100 552.0B params 1M context Proprietary

About this model

DeepSeek-V4.1-Flash is a multimodal mixture-of-experts language model designed for long-context, agentic, and vision-language workloads. It uses a 40-layer causal encoder-decoder (CED) Transformer with Compressed Sparse Attention 2 (CSA2), FP4 KV caching, Engram conditional memory, and DSpark speculative decoding, activating about 8B parameters during prefill and 16B during decode from a 552B-parameter backbone. The model supports up to one million tokens of context, native image-and-text input, continuously adjustable reasoning effort (1–100), and strong scores on terminal, software-engineering, and cybersecurity agent benchmarks when evaluated at maximum reasoning effort.

The open weights and reference inference code are released on Hugging Face under the MIT license, while the DeepSeek Batch API offers the same V4.1-Flash instruct capabilities for asynchronous, cost-efficient bulk inference. Post-training follows SFT, reinforcement learning, and on-policy distillation with large-scale synthetic agent data. On standard knowledge and reasoning suites the instruct model reaches competitive GPQA-Diamond, LiveBench, and AIME 2025 results; on agent harnesses it leads on Terminal-Bench 2.1, DeepSWE v1.1, CyberGym, and AutomationBench among compared DeepSeek variants.

DeepSeek-V4.1-Flash targets production agents that read long documents or repositories, run coding and terminal tools, and interpret charts or UI screenshots without sacrificing throughput. KV cache compression to roughly 890 bytes per token materially reduces memory for million-token sessions relative to earlier DeepSeek-Flash generations, making it practical for input-heavy pipelines served via API or self-hosted with multi-GPU clusters.

Benchmark Scores

BBH
86.1
HLE
39.1
DROP
87.9
MATH
61.1
MMLU
91.0
GSM8K
93.0
C-Eval
92.1
IFEval
89.5
AGIEval
83.4
DeepSWE
74.2
CyberGym
88.1
MMLU-Pro
81.2
MMMU-Pro
56.5
AIME-2025
87.5
HellaSwag
87.2
HumanEval
79.4
LiveBench
81.11
CodeForces
3471.0
SimpleBench
66.7
GPQA-Diamond
90.9
NL2Repo-Bench
65.4
AutomationBench
54.8
Agents-Last-Exam
31.8
Terminal-Bench-2.1
90.6
Terminal-Bench-3.0
30.0

Technical Specs

  • Parameters: 552.0B
  • Architecture: Mixture-of-Experts (MoE) Transformer
  • Context Window: 1,000,000 tokens
  • Input Modalities: text, image

Hardware Requirements

  • Compute: API only

Pricing

Input Output Currency
0.11 / 1M tokens 0.34 / 1M tokens USD