AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

562 models for "Vision" Compare
Molmo2-4B logo
Molmo2-4B
allenai

Molmo2 is a family of open vision-language models developed by the Allen Institute for AI (Ai2) that support image, video and multi-image understanding and grounding. Molmo2 models are trained on publicly available third party datasets as referenced in our technical report and Mo…

Multimodal 4.0B ↓ 47.9K
Florence-2-large-ft logo
Florence-2-large-ft
microsoft

Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks

Multimodal 0.77B ↓ 47.4K
instructblip-flan-t5-xl logo
instructblip-flan-t5-xl
Salesforce

InstructBLIP model using Flan-T5-xl as language model. InstructBLIP was introduced in the paper InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning by Dai et al.

Multimodal 4.0B ↓ 42.2K
InternVL3_5-8B logo
InternVL3_5-8B
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 8.5B ↓ 40.6K
GLM-4.5V logo
GLM-4.5V
zai-org

This model is part of the GLM-V family of models, introduced in the paper GLM-4.1V-Thinking and GLM-4.5V: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

Open Source ↓ 39.4K
Emu3-Chat-hf logo
Emu3-Chat-hf
BAAI

Emu3: Next-Token Prediction is All You Need

Multimodal 8.0B ↓ 38.1K
SmolVLM2-256M-Video-Instruct logo
SmolVLM2-256M-Video-Instruct
HuggingFaceTB

SmolVLM2-256M-Video is a lightweight multimodal model designed to analyze video content. The model processes videos, images, and text inputs to generate text outputs - whether answering questions about media files, comparing visual content, or transcribing text from images. Despi…

Multimodal 0.256B ↓ 38K
LFM2.5-VL-450M logo
LFM2.5-VL-450M
LiquidAI

LFM2.5‑VL-450M is Liquid AI's refreshed version of the first vision-language model, LFM2-VL-450M, built on an updated backbone LFM2.5-350M and tuned for stronger real-world performance. Find more about LFM2.5 family of models in our blog post.

Multimodal 0.45B ↓ 37K
MiniCPM-V-2_6 logo
MiniCPM-V-2_6
openbmb

Multimodal 8.0B ↓ 36.3K
InternVL3_5-1B logo
InternVL3_5-1B
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 1.1B ↓ 30K
MiniMax-VL-01 logo
MiniMax-VL-01
MiniMaxAI

1. Introduction We are delighted to introduce our MiniMax-VL-01 model. It adopts the "ViT-MLP-LLM" framework, which is a commonly used technique in the field of multimodal large language models. The model is initialized and trained with three key parts: a 303-million-parameter Vi…

Multimodal 456.0B ↓ 29.7K
ERNIE-4.5-VL-28B-A3B-PT logo
ERNIE-4.5-VL-28B-A3B-PT
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Multimodal 28.0B ↓ 29.2K