AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

562 models for "Vision" Compare
InternVL3-9B logo
InternVL3-9B
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 9.1B ↓ 4.3K
granite-3.1-1b-a400m-base logo
granite-3.1-1b-a400m-base
ibm-granite

Model Summary: Granite-3.1-1B-A400M-Base extends the context length of Granite-3.0-1B-A400M-Base from 4K to 128K using a progressive training strategy by increasing the supported context length in increments while adjusting RoPE theta until the model has successfully adapted to d…

Open Source 1.0B ↓ 4.2K
InternVL3-1B-Instruct logo
InternVL3-1B-Instruct
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 1.0B ↓ 4K
UI-TARS-2B-SFT logo
UI-TARS-2B-SFT
ByteDance-Seed

UI-TARS-2B-SFT UI-TARS-2B-SFT     UI-TARS-2B-gguf     UI-TARS-7B-SFT     UI-TARS-7B-DPO (Recommended)     UI-TARS-7B-gguf     UI-TARS-72B-SFT     UI-TARS-72B-DPO (Recommended) Introduction

Multimodal 2.0B ↓ 3.9K
OpenELM-270M logo
OpenELM-270M
apple

Sachin Mehta, Mohammad Hossein Sekhavat, Qingqing Cao, Maxwell Horton, Yanzi Jin, Chenfan Sun, Iman Mirzadeh, Mahyar Najibi, Dmitry Belenko, Peter Zatloukal, Mohammad Rastegari

Open Source 0.27B ↓ 3.8K
GLM-4.6V logo
GLM-4.6V
zai-org

This model is part of the GLM-V family of models, introduced in the paper GLM-4.1V-Thinking and GLM-4.5V: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

Open Source ↓ 3.8K
Olmo-Hybrid-Instruct-SFT-7B logo
Olmo-Hybrid-Instruct-SFT-7B
allenai

We expand on our Olmo model series by introducing Olmo Hybrid, a new 7B hybrid RNN model in the Olmo family. Olmo Hybrid dramatically outperforms Olmo 3 in final performance, consistently showing roughly 2x data efficiency on core evals over the course of our pretraining run. We…

Open Source 7.0B ↓ 3.7K
Molmo-7B-O-0924 logo
Molmo-7B-O-0924
allenai

Molmo is a family of open vision-language models developed by the Allen Institute for AI. Molmo models are trained on PixMo, a dataset of 1 million, highly-curated image-text pairs. It has state-of-the-art performance among multimodal models with a similar size while being fully…

Multimodal 7.7B ↓ 3.5K
InternVL2_5-8B-MPO logo
InternVL2_5-8B-MPO
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 8.1B ↓ 3.3K
MiniCPM-V-4_5-int4 logo
MiniCPM-V-4_5-int4
openbmb

A GPT-4o Level MLLM for Single Image, Multi Image and High-FPS Video Understanding on Your Phone

Multimodal 8.7B ↓ 3.3K
OpenELM-1_1B logo
OpenELM-1_1B
apple

Sachin Mehta, Mohammad Hossein Sekhavat, Qingqing Cao, Maxwell Horton, Yanzi Jin, Chenfan Sun, Iman Mirzadeh, Mahyar Najibi, Dmitry Belenko, Peter Zatloukal, Mohammad Rastegari

Open Source 1.0B ↓ 3K
InternVL3-38B-hf logo
InternVL3-38B-hf
OpenGVLab

InternVL3-38B Transformers 🤗 Implementation

Multimodal 38.4B ↓ 2.8K