AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

562 models for "Vision" Compare
Qwen2.5-VL-32B-Instruct logo
Qwen2.5-VL-32B-Instruct
Qwen

Latest Updates: In addition to the original formula, we have further enhanced Qwen2.5-VL-32B's mathematical and problem-solving abilities through reinforcement learning. This has also significantly improved the model's subjective user experience, with response styles adjusted to…

Multimodal 32.0B ↓ 791.1K
OpenELM-1_1B-Instruct logo
OpenELM-1_1B-Instruct
apple

Sachin Mehta, Mohammad Hossein Sekhavat, Qingqing Cao, Maxwell Horton, Yanzi Jin, Chenfan Sun, Iman Mirzadeh, Mahyar Najibi, Dmitry Belenko, Peter Zatloukal, Mohammad Rastegari

Open Source 1.1B ↓ 670.5K
Qwen2-VL-7B-Instruct logo
Qwen2-VL-7B-Instruct
Qwen

We're excited to unveil Qwen2-VL , the latest iteration of our Qwen-VL model, representing nearly a year of innovation.

Multimodal 7.0B ↓ 661.9K
InternVL2-1B logo
InternVL2-1B
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 Mini-InternVL\]](https://arxiv.org/abs/2410.16261) [\[📜 InternVL 2.5\]](https://huggingface.co/…

Multimodal 0.9B ↓ 648.9K
Cosmos-Reason2-2B logo
Cosmos-Reason2-2B
nvidia

Multimodal 2.0B ↓ 641.9K
blip2-opt-2.7b logo
blip2-opt-2.7b
Salesforce

BLIP-2 model, leveraging OPT-2.7b (a large language model with 2.7 billion parameters). It was introduced in the paper BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models by Li et al. and first released in this repository.

Open Source 2.7B ↓ 606.6K
InternVL2-2B logo
InternVL2-2B
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 Mini-InternVL\]](https://arxiv.org/abs/2410.16261) [\[📜 InternVL 2.5\]](https://huggingface.co/…

Multimodal 2.2B ↓ 561.7K
GOT-OCR2_0 logo
GOT-OCR2_0
stepfun-ai

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Multimodal 0.58B ↓ 546.1K
MiniMax-M2.5 logo
MiniMax-M2.5
MiniMaxAI

Join Our 💬 WeChat 🧩 Discord community. MiniMax Agent ⚡️ API MCP MiniMax Website 🤗 Hugging Face 🚀 Hugging Face API 🐙 GitHub 🤖️ ModelScope 📄 License: Modified-MIT

Open Source ↓ 532.9K
UI-TARS-1.5-7B logo
UI-TARS-1.5-7B
ByteDance-Seed

--- license: apache-2.0 language: - en pipeline tag: image-text-to-text tags: - multimodal - gui library name: transformers ---

Multimodal 7.0B ↓ 466K
Kimi-K2.6 logo
Kimi-K2.6
moonshotai

🤗   huggingchat     📰   Tech Blog

Open Source ↓ 446.1K
gemma-3-27b-it logo
gemma-3-27b-it
google

Multimodal 27.0B ↓ 417.6K