AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

562 models for "Vision" Compare
GLM-4.5V-FP8 logo
GLM-4.5V-FP8
zai-org

👋 Join our Discord communities. 📖 Check out the paper . 📍 Access the GLM-V series models via API on the ZhipuAI Open Platform .

Open Source ↓ 2.8K
FastVLM-1.5B logo
FastVLM-1.5B
apple

FastVLM: Efficient Vision Encoding for Vision Language Models

Multimodal 1.5B ↓ 2.7K
evo-1-8k-base logo
evo-1-8k-base
togethercomputer

We identified and fixed an issue related to a wrong permutation of some projections, which affects generation quality. To use the new model revision, please load as follows:

Open Source 7.0B ↓ 2.7K
MiniCPM-V-2_6-int4 logo
MiniCPM-V-2_6-int4
openbmb

[2025.01.14] 🔥🔥 We open source MiniCPM-o 2.6 , with significant performance improvement over MiniCPM-V 2.6 , and support real-time speech-to-speech conversation and multimodal live streaming. Try it now.

Multimodal 8.0B ↓ 2.7K
InternVL2_5-26B logo
InternVL2_5-26B
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 Mini-InternVL\]](https://arxiv.org/abs/2410.16261) [\[📜 InternVL 2.5\]](https://huggingface.co/…

Multimodal 26.0B ↓ 2.7K
InternVL3_5-38B-Instruct logo
InternVL3_5-38B-Instruct
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 38.4B ↓ 2.6K
LensVLM-9B logo
LensVLM-9B
apple

LensVLM is a 9B Vision Language Model (VLM) that scans compressed images of text, then selectively expands only the relevant pages to their uncompressed form via learned tools.

Multimodal 9.0B ↓ 2.6K
InternVL3_5-8B-Instruct logo
InternVL3_5-8B-Instruct
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 8.5B ↓ 2.4K
InternVL2_5-38B logo
InternVL2_5-38B
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 Mini-InternVL\]](https://arxiv.org/abs/2410.16261) [\[📜 InternVL 2.5\]](https://huggingface.co/…

Multimodal 38.4B ↓ 2.4K
LFM2-VL-3B logo
LFM2-VL-3B
LiquidAI

LFM2-VL-3B is the newest and most capable model in Liquid AI's multimodal LFM2-VL series, designed to process text and images with variable resolutions. Built on the LFM2 backbone, it extends the architecture for higher-capacity reasoning and stronger visual understanding while r…

Multimodal 3.0B ↓ 2.3K
InternVL3_5-8B-Flash logo
InternVL3_5-8B-Flash
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 8.0B ↓ 2.2K
OpenELM-3B logo
OpenELM-3B
apple

Sachin Mehta, Mohammad Hossein Sekhavat, Qingqing Cao, Maxwell Horton, Yanzi Jin, Chenfan Sun, Iman Mirzadeh, Mahyar Najibi, Dmitry Belenko, Peter Zatloukal, Mohammad Rastegari

Open Source 3.0B ↓ 2.2K