DeepSeek: DeepSeek V4 Flash Latest
About this model
DeepSeek V4 Flash is the efficient tier of the DeepSeek-V4 family: a 284-billion-parameter mixture-of-experts (MoE) language model with about 13 billion parameters activated per token. It pairs a hybrid attention stack (Compressed Sparse Attention plus Heavily Compressed Attention) with Manifold-Constrained Hyper-Connections and Muon optimization, targeting million-token contexts at a fraction of the compute and KV-cache footprint of earlier DeepSeek generations. Weights ship on Hugging Face in mixed FP4 and FP8 precision, with open inference recipes for multi-GPU serving and optional speculative decoding variants.
Post-training follows DeepSeek's two-stage recipe: domain-specific expert cultivation, then unified consolidation via on-policy distillation. The instruct model exposes Non-think, Think High, and Think Max reasoning modes; Max is where the strongest public benchmark numbers are reported. The July 2026 V4-Flash-0731 API checkpoint substantially raised agentic scores on terminal, repository, and cybersecurity-style suites while keeping the same structural footprint as the April preview release.
DeepSeek V4 Flash is aimed at long-horizon agents, software engineering, and competitive coding at lower serving cost than V4-Pro. It is competitive with much larger models on LiveCodeBench, SWE-Bench Verified, and Codeforces-style evaluation when run at high reasoning effort, and remains a practical open-weight option for teams that need very long context windows without provisioning a 1.6T-parameter Pro deployment.
Benchmark Scores
Technical Specs
- Parameters: 284.0B
- Architecture: Mixture-of-Experts Transformer
- Context Window: 1,000,000 tokens
- Input Modalities: text
Hardware Requirements
- VRAM: 192.0 GB
- Compute: 2x NVIDIA H200 141GB or single AMD MI300X 192GB
Pricing
| Input | Output | Currency |
|---|---|---|
| 0.00 / 1M tokens | 0.35 / 1M tokens | USD |