LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
Description: The NVIDIA GLM-5.2 NVFP4 model is the quantized version of ZAI’s GLM-5.2 model, which is an auto-regressive language model that uses an optimized transformer architecture. GLM-5.2 is a Mixture-of-Experts (MoE) model for reasoning and coding that uses sparse attention…
Description: The NVIDIA DeepSeek-V4-Pro-NVFP4 model is the quantized version of the DeepSeek-V4-Pro model, which is a Mixture-of-Experts (MoE) language model with 1.6 trillion total parameters and 49 billion activated parameters. For more information, please check here. The NVIDI…
Description: The NVIDIA DeepSeek-V4-Flash-NVFP4 model is a quantized version of DeepSeek AI's DeepSeek-V4-Flash model, an autoregressive Mixture-of-Experts language model that uses an optimized Transformer architecture with hybrid attention (Compressed Sparse Attention and Heavil…
Description MiniMax-M3 is a multimodal model with frontier-level coding and agentic capabilities, built on a Mixture-of-Experts architecture with a 1M-token context window. The model processes text, image, video, and computer use inputs and produces text outputs, with emphasis on…
Description: The NVIDIA Kimi-K2.7-Code NVFP4 model is the quantized version of the Moonshot AI's Kimi-K2.7-Code model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Kimi-K2.7-Code NV…
Description: The NVIDIA GLM-5.3-Flash NVFP4 model is the quantized version of ZAI's GLM-5.3-Flash model, which is an auto-regressive language model that uses an optimized transformer architecture. GLM-5.3-Flash is a natively multimodal Mixture-of-Experts (MoE) model for reasoning…
Description: The NVIDIA Qwen3.8-Flash-Next NVFP4 model is the quantized version of Alibaba's Qwen3.8-Flash-Next model, which is an auto-regressive language model that uses an optimized transformer architecture. Qwen3.8-Flash-Next is a causal language model with a vision encoder,…
Description: The NVIDIA Qwen3.6-27B NVFP4 model is the quantized version of Alibaba's Qwen3.6-27B model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Qwen3.6-27B NVFP4 model is quan…
Description: The NVIDIA Qwen3.8-27B NVFP4 model is a quantized version of Alibaba's Qwen3.8-27B model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information on the model, please check here. The model is quantized with Mod…
We introduce Olmo 3, a new family of 7B and 32B models both Instruct and Think variants. Long chain-of-thought thinking improves reasoning tasks like math and coding.
We introduce Olmo 3, a new family of 7B and 32B models both Instruct and Think variants. Long chain-of-thought thinking improves reasoning tasks like math and coding.
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Base-BF16