LLM Models · deepseek-ai
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
DeepSeek-V4-Pro-0813 is the official release of DeepSeek-V4-Pro , superseding the preview version, with greatly enhanced agentic capabilities and performance improvements that are especially pronounced in production environments. It is built on the DeepSeek-V4-Pro (Preview) model…
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash , superseding the preview version, with substantially enhanced agentic capabilities. It has the same model structure as DeepSeek-V4-Flash-DSpark, i.e. it comes with a speculative decoding module attached.
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
🌟 Github 📥 Model Download 📄 Paper Link 📄 Arxiv Paper Link DeepSeek-OCR: Contexts Optical Compression Explore the boundaries of visual-text compression.
We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With R…
DeepSeek-V3.2: Efficient Reasoning & Agentic AI
DeepSeek-V3-0324 demonstrates notable improvements over its predecessor, DeepSeek-V3, in several key aspects.
We present DeepSeek-V3, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token. To achieve efficient inference and cost-effective training, DeepSeek-V3 adopts Multi-head Latent Attention (MLA) and DeepSeekMoE architectures, w…
🌟 Github 📥 Model Download 📄 Paper Link 📄 Arxiv Paper Link DeepSeek-OCR 2: Visual Causal Flow Explore more human-like visual encoding.