LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
Hugging Face GitHub Launch Blog Documentation Technical Report License : Apache 2.0 Authors : Google DeepMind
Hugging Face GitHub Launch Blog Documentation Technical Report License : Apache 2.0 Authors : Google DeepMind
Hugging Face GitHub Launch Blog Documentation Technical Report License : Apache 2.0 Authors : Google DeepMind
Hugging Face GitHub Launch Blog Documentation Technical Report License : Apache 2.0 Authors : Google DeepMind
Hugging Face GitHub Launch Blog Documentation Technical Report License : Apache 2.0 Authors : Google DeepMind
:--- :--- Total Parameters 120B (12B active) Architecture LatentMoE - Mamba-2 + MoE + Attention hybrid with Multi-Token Prediction (MTP) Context Length Up to 1M tokens Minimum GPU Requirement 8× H100-80GB Supported Languages English, French, German, Italian, Japanese, Spanish, Ch…
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4
:--- :--- Total Parameters 120B (12B active) Architecture LatentMoE - Mamba-2 + MoE + Attention hybrid with Multi-Token Prediction (MTP) Context Length Up to 1M tokens Minimum GPU Requirement 1× B200 OR 1× DGX Spark Supported Languages English, French, German, Italian, Japanese,…
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
The NVIDIA Nemotron-3.5-Lightning-30B-A3B-NVFP4-DSpark model is the DSpark speculative decoding checkpoint for NVIDIA's Nemotron-3.5-Lightning-30B-A3B model family, which is a hybrid LatentMoE language model designed for reasoning, chat, and agentic workflows. For more informatio…
:--- :--- Total Parameters 120B (12B active) Architecture LatentMoE - Mamba-2 + MoE + Attention hybrid with Multi-Token Prediction (MTP) Context Length Up to 1M tokens Minimum GPU Requirement 2× H100-80GB Supported Languages English, French, German, Italian, Japanese, Spanish, Ch…
Try gpt-oss · Guides · Model card · OpenAI blog