AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

362 models for "Fine-tuned" Compare
GLM-5.2-NVFP4 logo
GLM-5.2-NVFP4
nvidia

Description: The NVIDIA GLM-5.2 NVFP4 model is the quantized version of ZAI’s GLM-5.2 model, which is an auto-regressive language model that uses an optimized transformer architecture. GLM-5.2 is a Mixture-of-Experts (MoE) model for reasoning and coding that uses sparse attention…

Open Source ★ 53.0 ↓ 450.6K
DeepSeek-V4-Pro-NVFP4 logo
DeepSeek-V4-Pro-NVFP4
nvidia

Description: The NVIDIA DeepSeek-V4-Pro-NVFP4 model is the quantized version of the DeepSeek-V4-Pro model, which is a Mixture-of-Experts (MoE) language model with 1.6 trillion total parameters and 49 billion activated parameters. For more information, please check here. The NVIDI…

Open Source ★ 53.0 ↓ 146.5K
DeepSeek-V4-Flash-NVFP4 logo
DeepSeek-V4-Flash-NVFP4
nvidia

Description: The NVIDIA DeepSeek-V4-Flash-NVFP4 model is a quantized version of DeepSeek AI's DeepSeek-V4-Flash model, an autoregressive Mixture-of-Experts language model that uses an optimized Transformer architecture with hybrid attention (Compressed Sparse Attention and Heavil…

Open Source ★ 52.0 ↓ 223.1K
MiniMax-M3-NVFP4 logo
MiniMax-M3-NVFP4
nvidia

Description MiniMax-M3 is a multimodal model with frontier-level coding and agentic capabilities, built on a Mixture-of-Experts architecture with a 1M-token context window. The model processes text, image, video, and computer use inputs and produces text outputs, with emphasis on…

Open Source ★ 45.0 ↓ 183K
Kimi-K2.7-Code-NVFP4 logo
Kimi-K2.7-Code-NVFP4
nvidia

Description: The NVIDIA Kimi-K2.7-Code NVFP4 model is the quantized version of the Moonshot AI's Kimi-K2.7-Code model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Kimi-K2.7-Code NV…

Code ★ 43.0 ↓ 294.8K
GLM-5.3-Flash-NVFP4 logo
GLM-5.3-Flash-NVFP4
nvidia

Description: The NVIDIA GLM-5.3-Flash NVFP4 model is the quantized version of ZAI's GLM-5.3-Flash model, which is an auto-regressive language model that uses an optimized transformer architecture. GLM-5.3-Flash is a natively multimodal Mixture-of-Experts (MoE) model for reasoning…

Open Source ★ 42.0 ↓ 87.5K
Qwen3.8-Flash-Next-NVFP4 logo
Qwen3.8-Flash-Next-NVFP4
nvidia

Description: The NVIDIA Qwen3.8-Flash-Next NVFP4 model is the quantized version of Alibaba's Qwen3.8-Flash-Next model, which is an auto-regressive language model that uses an optimized transformer architecture. Qwen3.8-Flash-Next is a causal language model with a vision encoder,…

Open Source ★ 40.0 ↓ 342.7K
Qwen3.6-27B-NVFP4 logo
Qwen3.6-27B-NVFP4
nvidia

Description: The NVIDIA Qwen3.6-27B NVFP4 model is the quantized version of Alibaba's Qwen3.6-27B model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Qwen3.6-27B NVFP4 model is quan…

Open Source 27.0B ★ 38.0 ↓ 614.8K
Qwen3.8-27B-NVFP4 logo
Qwen3.8-27B-NVFP4
nvidia

Description: The NVIDIA Qwen3.8-27B NVFP4 model is a quantized version of Alibaba's Qwen3.8-27B model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information on the model, please check here. The model is quantized with Mod…

Open Source 27.0B ★ 34.0 ↓ 622.9K
Olmo-3.1-7B-RL-Zero-Code logo
Olmo-3.1-7B-RL-Zero-Code
allenai

We introduce Olmo 3, a new family of 7B and 32B models both Instruct and Think variants. Long chain-of-thought thinking improves reasoning tasks like math and coding.

Code 7.0B ★ 31.0 ↓ 4.1K
Olmo-3-32B-Think-DPO logo
Olmo-3-32B-Think-DPO
allenai

We introduce Olmo 3, a new family of 7B and 32B models both Instruct and Think variants. Long chain-of-thought thinking improves reasoning tasks like math and coding.

Open Source 32.0B ★ 31.0 ↓ 4K
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Base-BF16 logo
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Base-BF16
nvidia

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Base-BF16

Open Source 30.0B ★ 24.0 ↓ 143.2K