LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
Hugging Face GitHub Launch Blog Documentation Technical Report License : Apache 2.0 Authors : Google DeepMind
:--- :--- Total Parameters 120B (12B active) Architecture LatentMoE - Mamba-2 + MoE + Attention hybrid with Multi-Token Prediction (MTP) Context Length Up to 1M tokens Minimum GPU Requirement 2× H100-80GB Supported Languages English, French, German, Italian, Japanese, Spanish, Ch…
🤗 HuggingFace 📔 Technical Report 📰 Blog Play around! 🗨️ Xiaomi MiMo Studio 🎨 Xiaomi MiMo API Platform
🤗 Hugging Face 🤖 ModelScope 🐙 OpenRouter
🤗 HuggingFace 📔 Technical Report 📰 Blog Play around! 🗨️ Xiaomi MiMo Studio 🎨 Xiaomi MiMo API Platform
Try gpt-oss · Guides · Model card · OpenAI blog
[!Note] This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.
[!Note] This repository contains FP8-quantized model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc. The quantization method is fin…
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
The NVIDIA Nemotron-3.5-Lightning-30B-A3B-NVFP4-DSpark model is the DSpark speculative decoding checkpoint for NVIDIA's Nemotron-3.5-Lightning-30B-A3B model family, which is a hybrid LatentMoE language model designed for reasoning, chat, and agentic workflows. For more informatio…
--- --- Developers Granite Team, IBM Model Type Decoder-only Dense Transformer (Reasoning) Architecture GraniteForCausalLM Base Model Granite-4.1-30B-Base Parameters 30B Context Length Natively Supports 128K (Long-context extension to 512K) Precision bfloat16 Tested Languages Eng…