AI Agent Hub
Back to models
DeepSeek: DeepSeek Flash Latest logo

DeepSeek: DeepSeek Flash Latest

Multimodal ~deepseek Released 2026-09-10
39.0 / 100 552.0B params 1M context Proprietary

About this model

DeepSeek Flash Latest (DeepSeek-V4.1-Flash) is DeepSeek's current default Flash-tier model on the API under the name deepseek-flash, superseding DeepSeek-V4-Flash and retired vision experiment routes. It is a native multimodal mixture-of-experts model with 552 billion backbone parameters, a one-million-token context window, and MIT-licensed open weights on Hugging Face.

The architecture combines a Causal Encoder-Decoder (CED) stack with Compressed Sparse Attention 2 (CSA2), aggressive KV-cache compression (about 890 bytes per token), Engram conditional memory, and DSpark speculative decoding. It activates roughly 8B parameters per token during prefill and 16B during decode, targeting input-heavy agent workloads with lower memory and bandwidth than prior V4-Flash generations while adding built-in image understanding via DeepSeek-ViT.

Post-training emphasizes controllable reasoning effort (1-100) and large-scale agentic RL. At maximum reasoning effort, official evaluations report strong agentic coding results (including Terminal-Bench 2.1 and DeepSWE v1.1) and competitive reasoning scores such as GPQA Diamond and an internal Codeforces rating, with downloadable weights and API access for production multimodal and coding-agent use.

Benchmark Scores

BBH
86.1
HLE
39.1
DROP
87.9
MATH
61.1
GSM8K
93.0
C-Eval
92.1
AGIEval
83.4
DeepSWE
74.2
CyberGym
88.1
MMLU-Pro
74.1
MMMU-Pro
56.5
SimpleQA
42.3
AIME-2026
100.0
HellaSwag
87.2
HumanEval
79.4
LiveBench
81.11
CodeForces
3471.0
SimpleBench
66.7
GPQA-Diamond
90.9
LiveCodeBench
81.11
NL2Repo-Bench
64.0
AutomationBench
54.8
Agents-Last-Exam
31.8
Creative-Writing
1540.1
Terminal-Bench-2.1
90.6
Terminal-Bench-3.0
30.0

Technical Specs

  • Parameters: 552.0B
  • Architecture: MoE Transformer (CED)
  • Context Window: 1,000,000 tokens
  • Input Modalities: text, image

Hardware Requirements

  • VRAM: 510.0 GB
  • Compute: Multi-GPU server (~512GB VRAM for FP8 weights)

Pricing

Input Output Currency
0.02 / 1M tokens 1.20 / 1M tokens USD