AI Agent Hub

LLM Models · nvidia

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

58 models from nvidia Compare
Qwen3.5-122B-A10B-NVFP4 logo
Qwen3.5-122B-A10B-NVFP4
nvidia

Description: The NVIDIA Qwen3.5-122B-A10B-NVFP4 model is the quantized version of Alibaba's Qwen3.5-122B-A10B model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Qwen3.5-122B-A10B N…

open-source 122.0B ★ 33.0 ↓ 840.3K
Qwen3.6-35B-A3B-NVFP4 logo
Qwen3.6-35B-A3B-NVFP4
nvidia

Description: The NVIDIA Qwen3.6-35B-A3B-NVFP4 model is the quantized version of Alibaba's Qwen3.6-35B-A3B model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Qwen3.6-35B-A3B-NVFP4 m…

open-source 35.0B ★ 32.0 ↓ 10.8M
Gemma-4-31B-IT-NVFP4 logo
Gemma-4-31B-IT-NVFP4
nvidia

Description: Gemma 4 31B IT is an open multimodal model built by Google DeepMind that handles text and image inputs, can process video as sequences of frames, and generates text output. It is designed to deliver frontier-level performance for reasoning, agentic workflows, coding,…

open-source 31.0B ★ 30.0 ↓ 2M
Gemma-4-26B-A4B-NVFP4 logo
Gemma-4-26B-A4B-NVFP4
nvidia

Description: Gemma 4 26B IT is an open multimodal model built by Google DeepMind that handles text and image inputs, can process video as sequences of frames, and generates text output. It is designed to deliver frontier-level performance for reasoning, agentic workflows, coding,…

open-source 26.0B ★ 26.0 ↓ 1.5M
NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 logo
NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4
nvidia

:--- :--- Total Parameters 120B (12B active) Architecture LatentMoE - Mamba-2 + MoE + Attention hybrid with Multi-Token Prediction (MTP) Context Length Up to 1M tokens Minimum GPU Requirement 1× B200 OR 1× DGX Spark Supported Languages English, French, German, Italian, Japanese,…

reasoning 120.0B ★ 26.0 ↓ 1.2M
NVIDIA-Nemotron-3-Super-120B-A12B-BF16 logo
NVIDIA-Nemotron-3-Super-120B-A12B-BF16
nvidia

:--- :--- Total Parameters 120B (12B active) Architecture LatentMoE - Mamba-2 + MoE + Attention hybrid with Multi-Token Prediction (MTP) Context Length Up to 1M tokens Minimum GPU Requirement 8× H100-80GB Supported Languages English, French, German, Italian, Japanese, Spanish, Ch…

reasoning 120.0B ★ 26.0 ↓ 1.1M
NVIDIA-Nemotron-3-Super-120B-A12B-FP8 logo
NVIDIA-Nemotron-3-Super-120B-A12B-FP8
nvidia

:--- :--- Total Parameters 120B (12B active) Architecture LatentMoE - Mamba-2 + MoE + Attention hybrid with Multi-Token Prediction (MTP) Context Length Up to 1M tokens Minimum GPU Requirement 2× H100-80GB Supported Languages English, French, German, Italian, Japanese, Spanish, Ch…

reasoning 120.0B ★ 26.0 ↓ 125.6K
NVIDIA: Nemotron 3 Super (free) logo
NVIDIA: Nemotron 3 Super (free)
nvidia

NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer...

closed-source ★ 26.0
NVIDIA: Nemotron 3 Super logo
NVIDIA: Nemotron 3 Super
nvidia

NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer...

closed-source ★ 26.0
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 logo
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4
nvidia

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4

open-source 30.0B ★ 24.0 ↓ 946.3K
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 logo
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
nvidia

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

open-source 30.0B ★ 24.0 ↓ 438.3K
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4-DSpark logo
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4-DSpark
nvidia

The NVIDIA Nemotron-3.5-Lightning-30B-A3B-NVFP4-DSpark model is the DSpark speculative decoding checkpoint for NVIDIA's Nemotron-3.5-Lightning-30B-A3B model family, which is a hybrid LatentMoE language model designed for reasoning, chat, and agentic workflows. For more informatio…

open-source 30.0B ★ 24.0 ↓ 200K