AI Agent Hub

LLM Models · nvidia

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

58 models from nvidia Compare
Kimi-K2.5-NVFP4 logo
Kimi-K2.5-NVFP4
nvidia

Description: The NVIDIA Kimi-K2.5-NVFP4 model is the quantized version of the Moonshot AI's Kimi-K2.5 model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Kimi-K2.5 NVFP4 model is qu…

open-source ↓ 156.1K
NVIDIA-Nemotron-Parse-v1.1 logo
NVIDIA-Nemotron-Parse-v1.1
nvidia

NVIDIA Nemotron Parse v1.1 is designed to understand document semantics and extract text and tables elements with spatial grounding. Given an image, NVIDIA Nemotron Parse v1.1 produces structured annotations, including formatted text, bounding-boxes and the corresponding semantic…

multimodal 0.885B ↓ 141K
Qwen3-8B-NVFP4 logo
Qwen3-8B-NVFP4
nvidia

Description: The NVIDIA Qwen3-8B FP4 model is the quantized version of Alibaba's Qwen3-8B model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Qwen3-8B FP4 model is quantized with Te…

open-source 8.0B ↓ 131.4K
LocateAnything-3B logo
LocateAnything-3B
nvidia

LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding

multimodal 3.0B ↓ 115.1K
Cosmos-Reason1-7B logo
Cosmos-Reason1-7B
nvidia

Cosmos-Reason1: Physical AI Common Sense and Embodied Reasoning Models

multimodal 7.0B ↓ 106.7K
Nemotron-Labs-Diffusion-8B-Base logo
Nemotron-Labs-Diffusion-8B-Base
nvidia

Nemotron-Labs-Diffusion is a tri-mode language model that supports both AR decoding and diffusion-based parallel decoding by simply switching the attention pattern of the same model during inference. The synergy between these two modes enables a third mode, called self-speculatio…

open-source 8.0B ↓ 95.5K
Qwen3-14B-NVFP4 logo
Qwen3-14B-NVFP4
nvidia

Description: The NVIDIA Qwen3-14B FP4 model is the quantized version of Alibaba's Qwen3-14B model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Qwen3-14B FP4 model is quantized with…

open-source 14.8B ↓ 92.1K
Llama-3.1-70B-Instruct-FP8 logo
Llama-3.1-70B-Instruct-FP8
nvidia

Description: The NVIDIA Llama 3.1 70B Instruct FP8 model is the quantized version of the Meta's Llama 3.1 70B Instruct model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Llama 3.1…

open-source 70.0B ↓ 88.6K
DeepSeek-R1-0528-NVFP4-v2 logo
DeepSeek-R1-0528-NVFP4-v2
nvidia

Description: The NVIDIA DeepSeek-R1-0528-FP4 v2 model is the quantized version of the DeepSeek AI's DeepSeek R1 0528 model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA DeepSeek R1…

reasoning ↓ 87.6K
NVIDIA: Nemotron 3.5 Content Safety (free) logo
NVIDIA: Nemotron 3.5 Content Safety (free)
nvidia

NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B. It moderates both inputs to and responses from LLMs and VLMs, accepting...

closed-source