AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

562 models for "Vision" Compare
POINTS-GUI-G logo
POINTS-GUI-G
tencent

- 🔜 Upcoming: The End-to-End GUI Agent Model is currently under active development and will be released in a subsequent update. Stay tuned! - 🚀 2026.02.06: We are happy to present POINTS-GUI-G , our specialized GUI Grounding Model. To facilitate reproducible evaluation, we provid…

Open Source ↓ 339
Penguin-VL-2B logo
Penguin-VL-2B
tencent

Penguin-VL Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders

Multimodal 2.0B ↓ 329
UI-Venus-Ground-7B logo
UI-Venus-Ground-7B
inclusionAI

UI-Venus This repository contains the UI-Venus model from the report UI-Venus Technical Report: Building High-performance UI Agents with RFT.

Open Source 7.0B ↓ 326
Intern-S2-397B-FP8 logo
Intern-S2-397B-FP8
internlm

💻Github Repo • 🤗HF Model Collections • 🤖ModelScope Collections • 💬Online Chat

Multimodal 397.0B ↓ 310
ZwZ-4B logo
ZwZ-4B
inclusionAI

ZwZ-4B is a fine-grained multimodal perception model built upon Qwen3-VL-4B. It is trained using Region-to-Image Distillation (R2I) combined with reinforcement learning, enabling superior fine-grained visual understanding in a single forward pass — no inference-time zooming or to…

Open Source 4.0B ↓ 309
Llama-SEA-LION-v2-8B-IT logo
Llama-SEA-LION-v2-8B-IT
aisingapore

Llama-SEA-LION-v2-8B-IT SEA-LION is a collection of Large Language Models (LLMs) which have been pretrained and instruct-tuned for the Southeast Asia (SEA) region.

Open Source 8.0B ↓ 300
Penguin-VL-8B logo
Penguin-VL-8B
tencent

Penguin-VL Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders

Multimodal 8.0B ↓ 298
Spatial-SSRL-7B logo
Spatial-SSRL-7B
internlm

📖 Paper 🏠 Github 🤗 Spatial-SSRL-7B Model 🤗 Spatial-SSRL-3B Model 🤗 Spatial-SSRL-Qwen3VL-4B Model 🤗 Spatial-SSRL-81k Dataset 📰 Daily Paper

Open Source 7.0B ↓ 297
PaCoRe-8B logo
PaCoRe-8B
stepfun-ai

PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning

Open Source 8.0B ↓ 281
RLVR-8B-0926 logo
RLVR-8B-0926
stepfun-ai

PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning

Open Source 8.0B ↓ 280
SAIL-7B logo
SAIL-7B
ByteDance-Seed

[\[📂 GitHub\]](https://github.com/bytedance/SAIL) [\[📜 paper\]](https://arxiv.org/abs/2504.10462) [\[🚀 Quick Start\]]( quick-start)

Open Source 7.0B ↓ 277
falcon-11B-vlm logo
falcon-11B-vlm
tiiuae

Falcon2-11B-vlm is an 11B parameters causal decoder-only model built by TII and trained on over 5,000B tokens of RefinedWeb enhanced with curated corpora. To bring vision capabilities, we integrate the pretrained CLIP ViT-L/14 vision encoder with our Falcon2-11B chat-finetuned mo…

Multimodal 11.0B ↓ 268