AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

440 models for "Base Model" Compare
GLM-5.3-Flash-NVFP4 logo
GLM-5.3-Flash-NVFP4
nvidia

Description: The NVIDIA GLM-5.3-Flash NVFP4 model is the quantized version of ZAI's GLM-5.3-Flash model, which is an auto-regressive language model that uses an optimized transformer architecture. GLM-5.3-Flash is a natively multimodal Mixture-of-Experts (MoE) model for reasoning…

Open Source ★ 42.0 ↓ 87.5K
GLM-5.3-Flash-BF16 logo
GLM-5.3-Flash-BF16
zai-org

👋 Join our WeChat or Discord community. 📖 Check out the GLM-5.3-Flash blog and GLM-5 Technical report . 📍 Use GLM-5.3-Flash API services on Z.ai API Platform.

Open Source ★ 42.0 ↓ 67.7K
Hy3-preview-Base logo
Hy3-preview-Base
tencent

           

Open Source ★ 42.0 ↓ 317
Qwen3.8-Flash-Next-NVFP4 logo
Qwen3.8-Flash-Next-NVFP4
nvidia

Description: The NVIDIA Qwen3.8-Flash-Next NVFP4 model is the quantized version of Alibaba's Qwen3.8-Flash-Next model, which is an auto-regressive language model that uses an optimized transformer architecture. Qwen3.8-Flash-Next is a causal language model with a vision encoder,…

Open Source ★ 40.0 ↓ 342.7K
DeepSeek-V4.1-Flash logo
DeepSeek-V4.1-Flash
deepseek-ai

DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

Open Source ★ 39.0 ↓ 1.3M
Qwen3.6-27B-NVFP4 logo
Qwen3.6-27B-NVFP4
nvidia

Description: The NVIDIA Qwen3.6-27B NVFP4 model is the quantized version of Alibaba's Qwen3.6-27B model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Qwen3.6-27B NVFP4 model is quan…

Open Source 27.0B ★ 38.0 ↓ 614.8K
DeepSeek-V4-Pro logo
DeepSeek-V4-Pro
deepseek-ai

DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence

Open Source ★ 36.0 ↓ 423.6K
DeepSeek-V4-Pro-DSpark logo
DeepSeek-V4-Pro-DSpark
deepseek-ai

DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence

Open Source ★ 36.0 ↓ 3.9K
DeepSeek-V4-Flash logo
DeepSeek-V4-Flash
deepseek-ai

DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence

Open Source ★ 35.0 ↓ 966.5K
Qwen3.8-27B-NVFP4 logo
Qwen3.8-27B-NVFP4
nvidia

Description: The NVIDIA Qwen3.8-27B NVFP4 model is a quantized version of Alibaba's Qwen3.8-27B model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information on the model, please check here. The model is quantized with Mod…

Open Source 27.0B ★ 34.0 ↓ 622.9K
Olmo-3.1-7B-RL-Zero-Code logo
Olmo-3.1-7B-RL-Zero-Code
allenai

We introduce Olmo 3, a new family of 7B and 32B models both Instruct and Think variants. Long chain-of-thought thinking improves reasoning tasks like math and coding.

Code 7.0B ★ 31.0 ↓ 4.1K
Olmo-3-32B-Think-DPO logo
Olmo-3-32B-Think-DPO
allenai

We introduce Olmo 3, a new family of 7B and 32B models both Instruct and Think variants. Long chain-of-thought thinking improves reasoning tasks like math and coding.

Open Source 32.0B ★ 31.0 ↓ 4K