AI Agent Hub

LLM Models · nvidia

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

58 models from nvidia Compare
GLM-5.2-NVFP4 logo
GLM-5.2-NVFP4
nvidia

Description: The NVIDIA GLM-5.2 NVFP4 model is the quantized version of ZAI’s GLM-5.2 model, which is an auto-regressive language model that uses an optimized transformer architecture. GLM-5.2 is a Mixture-of-Experts (MoE) model for reasoning and coding that uses sparse attention…

open-source ★ 53.0 ↓ 1.2M
DeepSeek-V4-Pro-NVFP4 logo
DeepSeek-V4-Pro-NVFP4
nvidia

Description: The NVIDIA DeepSeek-V4-Pro-NVFP4 model is the quantized version of the DeepSeek-V4-Pro model, which is a Mixture-of-Experts (MoE) language model with 1.6 trillion total parameters and 49 billion activated parameters. For more information, please check here. The NVIDI…

open-source ★ 53.0 ↓ 150.1K
DeepSeek-V4-Flash-NVFP4 logo
DeepSeek-V4-Flash-NVFP4
nvidia

Description: The NVIDIA DeepSeek-V4-Flash-NVFP4 model is a quantized version of DeepSeek AI's DeepSeek-V4-Flash model, an autoregressive Mixture-of-Experts language model that uses an optimized Transformer architecture with hybrid attention (Compressed Sparse Attention and Heavil…

open-source ★ 52.0 ↓ 253.4K
MiniMax-M3-NVFP4 logo
MiniMax-M3-NVFP4
nvidia

Description MiniMax-M3 is a multimodal model with frontier-level coding and agentic capabilities, built on a Mixture-of-Experts architecture with a 1M-token context window. The model processes text, image, video, and computer use inputs and produces text outputs, with emphasis on…

open-source ★ 45.0 ↓ 186.5K
Kimi-K2.7-Code-NVFP4 logo
Kimi-K2.7-Code-NVFP4
nvidia

Description: The NVIDIA Kimi-K2.7-Code NVFP4 model is the quantized version of the Moonshot AI's Kimi-K2.7-Code model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Kimi-K2.7-Code NV…

code ★ 43.0 ↓ 489.3K
Qwen3.6-27B-NVFP4 logo
Qwen3.6-27B-NVFP4
nvidia

Description: The NVIDIA Qwen3.6-27B NVFP4 model is the quantized version of Alibaba's Qwen3.6-27B model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Qwen3.6-27B NVFP4 model is quan…

open-source 27.0B ★ 38.0 ↓ 704K
NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4 logo
NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4
nvidia

:--- :--- Total Parameters 550B (55B active) Architecture LatentMoE - Mamba-2 + MoE + Attention hybrid with Multi-Token Prediction (MTP) Context Length Up to 1M tokens Minimum GPU Requirement 4xGB200, 4xB200, 4x GB300, 4x B300, 8xH100 Supported Languages English, French, Spanish,…

open-source 550.0B ★ 38.0 ↓ 416K
NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 logo
NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
nvidia

:--- :--- Total Parameters 550B (55B active) Architecture LatentMoE - Mamba-2 + MoE + Attention hybrid with Multi-Token Prediction (MTP) Context Length Up to 1M tokens Minimum GPU Requirement 8x GB200/B200/GB300/B300, 16x H100, 8x H200 Supported Languages English, French, Spanish…

open-source 550.0B ★ 38.0 ↓ 296.5K
NVIDIA: Nemotron 3 Ultra (free) logo
NVIDIA: Nemotron 3 Ultra (free)
nvidia

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

closed-source ★ 38.0
NVIDIA: Nemotron 3 Ultra (batch) logo
NVIDIA: Nemotron 3 Ultra (batch)
nvidia

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

closed-source ★ 38.0
NVIDIA: Nemotron 3 Ultra logo
NVIDIA: Nemotron 3 Ultra
nvidia

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

closed-source ★ 38.0
Qwen3.5-397B-A17B-NVFP4 logo
Qwen3.5-397B-A17B-NVFP4
nvidia

Description: The NVIDIA Qwen3.5-397B-A17B NVFP4 model is the quantized version of Alibaba's Qwen3.5-397B-A17B model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Qwen3.5-397B-A17B N…

open-source 397.0B ★ 34.0 ↓ 186K