AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

562 models for "Vision" Compare
GLM-5.3 logo
GLM-5.3
zai-org

GLM-5.3 uses the same base model as GLM-5.2 — every gain comes from post-training. Compared with GLM-5.2, it is much better at complex coding and long-horizon tasks:

Open Source ★ 45.0 ↓ 1.5M
MiniMax-M3-NVFP4 logo
MiniMax-M3-NVFP4
nvidia

Description MiniMax-M3 is a multimodal model with frontier-level coding and agentic capabilities, built on a Mixture-of-Experts architecture with a 1M-token context window. The model processes text, image, video, and computer use inputs and produces text outputs, with emphasis on…

Open Source ★ 45.0 ↓ 183K
GLM-5.3-BF16 logo
GLM-5.3-BF16
zai-org

GLM-5.3 uses the same base model as GLM-5.2 — every gain comes from post-training. Compared with GLM-5.2, it is much better at complex coding and long-horizon tasks:

Open Source ★ 45.0 ↓ 38.9K
Qwen: Qwen3.8 Max (0902) logo
Qwen: Qwen3.8 Max (0902)
qwen

Qwen3.8 Max 0902 is an updated snapshot of Qwen3.8 Max from Alibaba's Qwen team. It is a 2.4-trillion-parameter mixture-of-experts model that accepts text, image, and video input and returns text,...

Closed Source 2400.0B ★ 45.0
Kimi-K3 logo
Kimi-K3
moonshotai

📰   Tech Blog     📄   Full Report

Open Source ★ 44.0 ↓ 1.2M
GLM-5.3-Flash logo
GLM-5.3-Flash
zai-org

👋 Join our WeChat or Discord community. 📖 Check out the GLM-5.3-Flash blog and GLM-5 Technical report . 📍 Use GLM-5.3-Flash API services on Z.ai API Platform.

Open Source ★ 42.0 ↓ 6.4M
GLM-5.3-Flash-NVFP4 logo
GLM-5.3-Flash-NVFP4
nvidia

Description: The NVIDIA GLM-5.3-Flash NVFP4 model is the quantized version of ZAI's GLM-5.3-Flash model, which is an auto-regressive language model that uses an optimized transformer architecture. GLM-5.3-Flash is a natively multimodal Mixture-of-Experts (MoE) model for reasoning…

Open Source ★ 42.0 ↓ 87.5K
GLM-5.3-Flash-BF16 logo
GLM-5.3-Flash-BF16
zai-org

👋 Join our WeChat or Discord community. 📖 Check out the GLM-5.3-Flash blog and GLM-5 Technical report . 📍 Use GLM-5.3-Flash API services on Z.ai API Platform.

Open Source ★ 42.0 ↓ 67.7K
Qwen3.8-Flash-Next logo
Qwen3.8-Flash-Next
Qwen

[!Note] This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, TokenSpeed, etc.

Open Source ★ 40.0 ↓ 1.6M
Qwen3.8-Flash-Next-NVFP4 logo
Qwen3.8-Flash-Next-NVFP4
nvidia

Description: The NVIDIA Qwen3.8-Flash-Next NVFP4 model is the quantized version of Alibaba's Qwen3.8-Flash-Next model, which is an auto-regressive language model that uses an optimized transformer architecture. Qwen3.8-Flash-Next is a causal language model with a vision encoder,…

Open Source ★ 40.0 ↓ 342.7K
DeepSeek-V4.1-Flash logo
DeepSeek-V4.1-Flash
deepseek-ai

DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

Open Source ★ 39.0 ↓ 1.3M
DeepSeek: DeepSeek V4.1 Flash (batch) logo
DeepSeek: DeepSeek V4.1 Flash (batch)
deepseek

DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on...

Multimodal 552.0B ★ 39.0