AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

562 models for "Vision" Compare
OpenELM-270M-Instruct logo
OpenELM-270M-Instruct
apple

Sachin Mehta, Mohammad Hossein Sekhavat, Qingqing Cao, Maxwell Horton, Yanzi Jin, Chenfan Sun, Iman Mirzadeh, Mahyar Najibi, Dmitry Belenko, Peter Zatloukal, Mohammad Rastegari

Open Source 0.27B ↓ 1.5K
MiMo-VL-7B-RL-2508 logo
MiMo-VL-7B-RL-2508
XiaomiMiMo

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ MiMo-VL Technical Report ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

Multimodal 7.0B ↓ 1.5K
UI-TARS-7B-DPO logo
UI-TARS-7B-DPO
ByteDance-Seed

UI-TARS-7B-DPO UI-TARS-2B-SFT     UI-TARS-7B-SFT     UI-TARS-7B-DPO (Recommended)     UI-TARS-72B-SFT     UI-TARS-72B-DPO (Recommended) Introduction

Multimodal 7.0B ↓ 1.4K
InternVL2-26B logo
InternVL2-26B
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 Mini-InternVL\]](https://arxiv.org/abs/2410.16261) [\[📜 InternVL 2.5\]](https://huggingface.co/…

Multimodal 25.5B ↓ 1.4K
InternVL3_5-30B-A3B-HF logo
InternVL3_5-30B-A3B-HF
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 30.0B ↓ 1.4K
UI-TARS-7B-SFT logo
UI-TARS-7B-SFT
ByteDance-Seed

UI-TARS-7B-SFT UI-TARS-2B-SFT     UI-TARS-7B-SFT     UI-TARS-7B-DPO (Recommended)     UI-TARS-72B-SFT     UI-TARS-72B-DPO (Recommended) Introduction

Open Source 7.0B ↓ 1.3K
UI-Venus-1.5-8B logo
UI-Venus-1.5-8B
inclusionAI

UI-Venus-1.5 model This repository contains the UI-Venus model from the report UI-Venus-1.5 Technical Report. UI-Venus 1.5 is a unified, end-to-end GUI Agent designed for robust real-world applications. The model family includes two dense (2B/8B) and one MoE (30B-A3B) variants to…

Multimodal 8.0B ↓ 1.3K
FastVLM-7B logo
FastVLM-7B
apple

FastVLM: Efficient Vision Encoding for Vision Language Models

Multimodal 7.0B ↓ 1.2K
Gemma-SEA-LION-v4-4B-VL logo
Gemma-SEA-LION-v4-4B-VL
aisingapore

SEA-LION is a collection of Large Language Models (LLMs) which have been pretrained and instruct-tuned for the Southeast Asia (SEA) region.

Multimodal 4.0B ↓ 1.2K
MiMo-VL-7B-RL logo
MiMo-VL-7B-RL
XiaomiMiMo

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ MiMo-VL Technical Report ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

Multimodal 7.0B ↓ 1.2K
Gemma-SEA-LION-v4-27B-IT logo
Gemma-SEA-LION-v4-27B-IT
aisingapore

SEA-LION is a collection of Large Language Models (LLMs) which have been pretrained and instruct-tuned for the Southeast Asia (SEA) region.

Open Source 27.0B ↓ 1.2K
Gemma-SEA-LION-v3-9B-IT logo
Gemma-SEA-LION-v3-9B-IT
aisingapore

SEA-LION is a collection of Large Language Models (LLMs) which have been pretrained and instruct-tuned for the Southeast Asia (SEA) region.

Open Source 9.0B ↓ 1.2K