AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

562 models for "Vision" Compare
InternVL2-8B-MPO logo
InternVL2-8B-MPO
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL/tree/main/internvl chat/shell/internvl2.0 mpo) [\[🆕 Blog\]](https://internvl.github.io/blog/2024-11-14-InternVL-2.0-MPO/) [\[📜 Paper\]](https://arxiv.org/abs/2411.10442) [\[📖 Documents\]](https://internvl.readthedocs.io/en/late…

Multimodal 8.1B ↓ 476
xgen-mm-phi3-mini-instruct-r-v1 logo
xgen-mm-phi3-mini-instruct-r-v1
Salesforce

📣 News 📌 [08/19/2024] xGen-MM-v1.5 released: - 🤗 xgen-mm-phi3-mini-instruct-interleave-r-v1.5 - 🤗 xgen-mm-phi3-mini-base-r-v1.5 - 🤗 xgen-mm-phi3-mini-instruct-singleimg-r-v1.5 - 🤗 xgen-mm-phi3-mini-instruct-dpo-r-v1.5

Multimodal 5.0B ↓ 472
MiniCPM-V-4_5-GPTQ logo
MiniCPM-V-4_5-GPTQ
openbmb

A GPT-4o Level MLLM for Single Image, Multi Image and High-FPS Video Understanding on Your Phone

Open Source ↓ 462
Intern-S2-Preview-397B logo
Intern-S2-Preview-397B
internlm

💻Github Repo • 🤗HF Model Collections • 🤖ModelScope Collections • 💬Online Chat

Open Source 397.0B ↓ 445
Emu3-Chat logo
Emu3-Chat
BAAI

Emu3: Next-Token Prediction is All You Need

Open Source ↓ 434
internlm-xcomposer2-vl-7b-4bit logo
internlm-xcomposer2-vl-7b-4bit
internlm

InternLM-XComposer2 is a vision-language large model (VLLM) based on InternLM2 for advanced text-image comprehension and composition.

Multimodal 7.0B ↓ 423
Yi-VL-6B logo
Yi-VL-6B
01-ai

🤗 Hugging Face • 🤖 ModelScope • 🟣 wisemodel

Multimodal 6.0B ↓ 421
UI-Venus-1.5-30B-A3B logo
UI-Venus-1.5-30B-A3B
inclusionAI

UI-Venus-1.5 model This repository contains the UI-Venus model from the report UI-Venus-1.5 Technical Report. UI-Venus 1.5 is a unified, end-to-end GUI Agent designed for robust real-world applications. The model family includes two dense (2B/8B) and one MoE (30B-A3B) variants to…

Multimodal 31.0B ↓ 420
CapRL-Qwen3VL-2B logo
CapRL-Qwen3VL-2B
internlm

CapRL 📖 Paper 🏠 Github 🤗 CapRL Collection 🤗 Daily Paper

Multimodal 2.0B ↓ 399
Yi-VL-34B logo
Yi-VL-34B
01-ai

🤗 Hugging Face • 🤖 ModelScope • 🟣 wisemodel

Multimodal 34.0B ↓ 396
Qianfan-VL-8B logo
Qianfan-VL-8B
baidu

Qianfan-VL: Domain-Enhanced Universal Vision-Language Models

Multimodal 8.0B ↓ 395
AREX-Turbo logo
AREX-Turbo
BAAI

Towards a Recursively Self-Improving Agent for Deep Research

Open Source ↓ 380