AI Agent Hub

LLM Models · nvidia

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

58 models from nvidia Compare
NVIDIA-Nemotron-Nano-12B-v2-VL-FP8 logo
NVIDIA-Nemotron-Nano-12B-v2-VL-FP8
nvidia

NVIDIA-Nemotron-Nano-VL-12B-V2-FP8 is the quantized version of the NVIDIA Nemotron Nano VL V2 model, which is an auto-regressive vision language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Nemotron Nano VL FP4 QAD mod…

multimodal 12.0B ★ 9.0 ↓ 125.7K
NVIDIA-Nemotron-Nano-9B-v2-FP8 logo
NVIDIA-Nemotron-Nano-9B-v2-FP8
nvidia

The pretraining data has a cutoff date of September 2024.

open-source 9.0B ★ 9.0 ↓ 121.3K
Cosmos-Reason2-2B logo
Cosmos-Reason2-2B
nvidia

multimodal 2.0B ↓ 982.2K
MiniMax-M2.5-NVFP4 logo
MiniMax-M2.5-NVFP4
nvidia

Description: The NVIDIA MiniMax-M2.5-NVFP4 model is the quantized version of MiniMax's MiniMax-M2.5 model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA MiniMax-M2.5 NVFP4 model is q…

open-source ↓ 776.5K
Cosmos-Reason2-8B logo
Cosmos-Reason2-8B
nvidia

multimodal 8.0B ↓ 527K
Qwen3-8B-FP8 logo
Qwen3-8B-FP8
nvidia

Description: The NVIDIA Qwen3-8B FP8 model is the quantized version of Alibaba's Qwen3-8B model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Qwen3-8B FP8 model is quantized with Te…

open-source 8.0B ↓ 437.5K
NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-NVFP4 logo
NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-NVFP4
nvidia

NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-NVFP4

open-source 75.0B ↓ 437.2K
Llama-3.1-Nemotron-Nano-VL-8B-V1 logo
Llama-3.1-Nemotron-Nano-VL-8B-V1
nvidia

Llama Nemotron Nano VL is a leading document intelligence vision language model (VLMs) that enables the ability to query and summarize images from the physical or virtual world. Llama Nemotron Nano VL is deployable in the data center, cloud and at the edge, including Jetson Orin…

multimodal 9.0B ↓ 301.4K
Llama-3.1-Nemotron-Nano-8B-v1 logo
Llama-3.1-Nemotron-Nano-8B-v1
nvidia

Llama-3.1-Nemotron-Nano-8B-v1 is a large language model (LLM) which is a derivative of Meta Llama-3.1-8B-Instruct (AKA the reference model). It is a reasoning model that is post trained for reasoning, human chat preferences, and tasks, such as RAG and tool calling.

reasoning 8.0B ↓ 218.1K
Llama-3_3-Nemotron-Super-49B-v1 logo
Llama-3_3-Nemotron-Super-49B-v1
nvidia

Llama-3.3-Nemotron-Super-49B-v1 is a large language model (LLM) which is a derivative of Meta Llama-3.3-70B-Instruct (AKA the reference model ). It is a reasoning model that is post trained for reasoning, human chat preferences, and tasks, such as RAG and tool calling. The model…

open-source 49.0B ↓ 176.5K
Kimi-K2.6-NVFP4 logo
Kimi-K2.6-NVFP4
nvidia

Description: The NVIDIA Kimi-K2.6-NVFP4 model is the quantized version of the Moonshot AI's Kimi-K2.6 model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Kimi-K2.6 NVFP4 model is qu…

open-source ↓ 162.3K
Nemotron-Labs-Diffusion-8B logo
Nemotron-Labs-Diffusion-8B
nvidia

Nemotron-Labs-Diffusion is a tri-mode language model that supports both AR decoding and diffusion-based parallel decoding by simply switching the attention pattern of the same model during inference. The synergy between these two modes enables a third mode, called self-speculatio…

open-source 8.0B ↓ 160.2K