LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
Description: The NVIDIA DeepSeek-V4-Flash-NVFP4 model is a quantized version of DeepSeek AI's DeepSeek-V4-Flash model, an autoregressive Mixture-of-Experts language model that uses an optimized Transformer architecture with hybrid attention (Compressed Sparse Attention and Heavil…
Description: The NVIDIA Kimi-K2.7-Code NVFP4 model is the quantized version of the Moonshot AI's Kimi-K2.7-Code model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Kimi-K2.7-Code NV…
Kimi K2.7 Code is a coding-focused agentic model built upon Kimi K2.6. With substantial improvements on real-world long-horizon coding tasks, it strengthens end-to-end task completion across complex software engineering workflows while improving token efficiency, reducing thinkin…
🤗 HuggingFace 📰 Blog 🎨 Xiaomi MiMo API Platform 🗨️ Xiaomi MiMo Studio
🎨 Xiaomi MiMo API Platform (Request Access) 🗨️ Xiaomi MiMo Studio (Free Trial)
Kimi K2 Instruct is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32 billion active per forward pass. It is optimized for...
:--- :--- Total Parameters 550B (55B active) Architecture LatentMoE - Mamba-2 + MoE + Attention hybrid with Multi-Token Prediction (MTP) Context Length Up to 1M tokens Minimum GPU Requirement 4xGB200, 4xB200, 4x GB300, 4x B300, 8xH100 Supported Languages English, French, Spanish,…
🤗 HuggingFace 📰 Blog 🎨 Xiaomi MiMo API Platform 🗨️ Xiaomi MiMo Studio