LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
Description: The NVIDIA Kimi-K2-Thinking-NVFP4 model is the quantized version of the Moonshot AI's Kimi-K2-Thinking model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Kimi-K2-Think…
Description: The NVIDIA Qwen3-Next-80B-A3B-Instruct NVFP4 model is the quantized version of Alibaba's Qwen3-Next-80B-A3B-Instruct model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA…
Try gpt-oss · Guides · Model card · OpenAI blog
--- pipeline tag: text-generation base model: google/diffusiongemma-26B-A4B-it license: apache-2.0 license name: apache-license-2.0 license link: https://ai.google.dev/gemma/apache 2 tags: - nvidia - ModelOpt - DiffusionGemma-26B-A4B-IT - quantized - NVFP4 - nvfp4 ---
The post-training data has a cutoff date of November 28, 2025\. The pre-training data has a cutoff date of June 25, 2025\.
The post-training data has a cutoff date of November 28, 2025\. The pre-training data has a cutoff date of June 25, 2025\.
The post-training data has a cutoff date of November 28, 2025\. The pre-training data has a cutoff date of June 25, 2025\.
NVIDIA-Nemotron-3-Nano-30B-A3B-Base-BF16
[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.
[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.
[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.
The advanced capabilities of the ERNIE 4.5 models, particularly the MoE-based A47B and A3B series, are underpinned by several key technical innovations: