LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
:--- :--- Total Parameters 550B (55B active) Architecture LatentMoE - Mamba-2 + MoE + Attention hybrid with Multi-Token Prediction (MTP) Context Length Up to 1M tokens Minimum GPU Requirement 8x GB200/B200/GB300/B300, 16x H100, 8x H200 Supported Languages English, French, Spanish…
🤗 Hugging Face 🤖 ModelScope 🐙 OpenRouter
🤗 Hugging Face 🤖 ModelScope 🐙 OpenRouter
🤗 Hugging Face 🤖 ModelScope 🐙 OpenRouter
🤗 Hugging Face 🤖 ModelScope 🐙 OpenRouter
🤗 Hugging Face 🤖 ModelScope 🐙 OpenRouter
🤗 Hugging Face 🤖 ModelScope 🐙 OpenRouter
🤗 Hugging Face 🤖 ModelScope 🐙 OpenRouter
🤗 HuggingFace 📰 Blog 🎨 Xiaomi MiMo API Platform 🗨️ Xiaomi MiMo Studio
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...