AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

562 models for "Vision" Compare
ERNIE-4.5-VL-28B-A3B-Base-PT logo
ERNIE-4.5-VL-28B-A3B-Base-PT
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Multimodal 28.0B ↓ 212
Qianfan-VL-70B logo
Qianfan-VL-70B
baidu

Qianfan-VL: Domain-Enhanced Universal Vision-Language Models

Multimodal 70.0B ↓ 208
UI-TARS-72B-SFT logo
UI-TARS-72B-SFT
ByteDance-Seed

UI-TARS-72B-SFT UI-TARS-2B-SFT     UI-TARS-7B-SFT     UI-TARS-7B-DPO (Recommended)     UI-TARS-72B-SFT     UI-TARS-72B-DPO (Recommended) Introduction

Open Source 72.0B ↓ 204
Qwen-SEA-LION-v4-32B-IT-4BIT logo
Qwen-SEA-LION-v4-32B-IT-4BIT
aisingapore

Qwen-SEA-LION-v4-32B-IT-4BIT (GPTQ model)

Open Source 32.0B ↓ 204
Intern-S1-mini-FP8 logo
Intern-S1-mini-FP8
internlm

💻Github Repo • 🤗Model Collections • 📜Technical Report • 💬Online Chat

Open Source ↓ 204
GTA1-7B logo
GTA1-7B
Salesforce

Reinforcement learning (RL) (e.g., GRPO) helps with grounding because of its inherent objective alignment—rewarding successful clicks—rather than encouraging long textual Chain-of-Thought (CoT) reasoning. Unlike approaches that rely heavily on verbose CoT reasoning, GRPO directly…

Open Source 7.0B ↓ 204
SmolVLM-Synthetic logo
SmolVLM-Synthetic
HuggingFaceTB

SmolVLM is a compact open multimodal model that accepts arbitrary sequences of image and text inputs to produce text outputs. Designed for efficiency, SmolVLM can answer questions about images, describe visual content, create stories grounded on multiple images, or function as a…

Multimodal ↓ 203
Qianfan-VL-3B logo
Qianfan-VL-3B
baidu

Qianfan-VL: Domain-Enhanced Universal Vision-Language Models

Multimodal 3.0B ↓ 202
ERNIE-4.5-21B-A3B-Paddle logo
ERNIE-4.5-21B-A3B-Paddle
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Open Source 21.0B ↓ 194
ZwZ-8B logo
ZwZ-8B
inclusionAI

ZwZ-8B is a fine-grained multimodal perception model built upon Qwen3-VL-8B. It is trained using Region-to-Image Distillation (R2I) combined with reinforcement learning, enabling superior fine-grained visual understanding in a single forward pass — no inference-time zooming or to…

Open Source 8.0B ↓ 191
Llama-3.1-8B-Dragonfly-v2 logo
Llama-3.1-8B-Dragonfly-v2
togethercomputer

Note: Users are permitted to use this model in accordance with the Llama 3.1 Community License Agreement.

Open Source 8.0B ↓ 184
sarashina1-7b logo
sarashina1-7b
sbintuitions

This repository provides Japanese language models trained by SB Intuitions.

Open Source 7.0B ↓ 178