AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

562 models for "Vision" Compare
VISTA-9B logo
VISTA-9B
inclusionAI

VISTA-9B are GUI-grounding vision-language models trained from Qwen3.5 9B backbones with VISTA: View-Consistent Self-Verified Training for GUI Grounding .

Open Source 9.0B ↓ 637
Falcon-E-1B-Instruct logo
Falcon-E-1B-Instruct
tiiuae

0. TL;DR 1. Model Details 2. Training Details 3. Usage 4. Evaluation 5. Citation

Open Source 1.0B ↓ 615
ERNIE-4.5-21B-A3B-Base-PT logo
ERNIE-4.5-21B-A3B-Base-PT
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Open Source 21.0B ↓ 611
SmolVLM-Base logo
SmolVLM-Base
HuggingFaceTB

SmolVLM is a compact open multimodal model that accepts arbitrary sequences of image and text inputs to produce text outputs. Designed for efficiency, SmolVLM can answer questions about images, describe visual content, create stories grounded on multiple images, or function as a…

Multimodal 2.2B ↓ 583
xgen-mm-phi3-mini-instruct-interleave-r-v1.5 logo
xgen-mm-phi3-mini-instruct-interleave-r-v1.5
Salesforce

Model description xGen-MM is a series of the latest foundational Large Multimodal Models (LMMs) developed by Salesforce AI Research. This series advances upon the successful designs of the BLIP series, incorporating fundamental enhancements that ensure a more robust and superior…

Open Source ↓ 573
OpenELM-450M logo
OpenELM-450M
apple

Sachin Mehta, Mohammad Hossein Sekhavat, Qingqing Cao, Maxwell Horton, Yanzi Jin, Chenfan Sun, Iman Mirzadeh, Mahyar Najibi, Dmitry Belenko, Peter Zatloukal, Mohammad Rastegari

Open Source ↓ 556
InternVL3_5-14B-Instruct logo
InternVL3_5-14B-Instruct
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 15.1B ↓ 555
InternVL3_5-38B-Flash logo
InternVL3_5-38B-Flash
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 38.0B ↓ 537
Emu3-Gen-hf logo
Emu3-Gen-hf
BAAI

Emu3: Next-Token Prediction is All You Need

Open Source ↓ 524
cogagent-9b-20241220 logo
cogagent-9b-20241220
zai-org

🌐 Github 🤗 Huggingface Space 📄 Technical Report 📜 arxiv paper

Multimodal 9.0B ↓ 496
InternVL3_5-2B-Flash logo
InternVL3_5-2B-Flash
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 2.3B ↓ 488
MiMo-VL-7B-SFT logo
MiMo-VL-7B-SFT
XiaomiMiMo

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ MiMo-VL Technical Report ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

Multimodal 7.0B ↓ 480