AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

562 models for "Vision" Compare
InternVL2-4B logo
InternVL2-4B
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 Mini-InternVL\]](https://arxiv.org/abs/2410.16261) [\[📜 InternVL 2.5\]](https://huggingface.co/…

Multimodal 4.2B ↓ 14.5K
InternVL3_5-1B-HF logo
InternVL3_5-1B-HF
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 1.1B ↓ 14.4K
GOT-OCR-2.0-hf logo
GOT-OCR-2.0-hf
stepfun-ai

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model - HF Transformers 🤗 implementation

Multimodal 0.6B ↓ 14.3K
Kimi-VL-A3B-Thinking logo
Kimi-VL-A3B-Thinking
moonshotai

[!Warning] This model has a new version: Kimi-VL-A3B-Thinking-2506. Please consider using the new 2506 version for better abilties on general visual understanding, reasoning, video and agent scenarios.

Multimodal 16.0B ↓ 14K
LFM2.5-1.2B-Base logo
LFM2.5-1.2B-Base
LiquidAI

LFM2.5 is a new family of hybrid models designed for on-device deployment . It builds on the LFM2 architecture with extended pre-training and reinforcement learning.

Open Source 1.2B ↓ 13.8K
MiniCPM-V-4_5-AWQ logo
MiniCPM-V-4_5-AWQ
openbmb

A GPT-4o Level MLLM for Single Image, Multi Image and Video Understanding on Your Phone

Multimodal 8.7B ↓ 13.7K
InternVL-Chat-V1-5 logo
InternVL-Chat-V1-5
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 Mini-InternVL\]](https://arxiv.org/abs/2410.16261) [\[📜 InternVL 2.5\]](https://huggingface.co/…

Multimodal 25.5B ↓ 13.6K
LFM2-VL-450M logo
LFM2-VL-450M
LiquidAI

LFM2‑VL is Liquid AI's first series of multimodal models, designed to process text and images with variable resolutions. Built on the LFM2 backbone, it is optimized for low-latency and edge AI applications.

Multimodal 0.45B ↓ 13.3K
MiniMax-M2.1 logo
MiniMax-M2.1
MiniMaxAI

Join Our 💬 WeChat 🧩 Discord community. MiniMax Agent ⚡️ API MCP MiniMax Website 🤗 Hugging Face 🐙 GitHub 🤖️ ModelScope 📄 License: Modified-MIT

Open Source ↓ 12.5K
InternVL-Chat-V1-2 logo
InternVL-Chat-V1-2
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 Mini-InternVL\]](https://arxiv.org/abs/2410.16261) [\[📜 InternVL 2.5\]](https://huggingface.co/…

Multimodal 40.0B ↓ 12.4K
UI-Venus-1.5-2B logo
UI-Venus-1.5-2B
inclusionAI

UI-Venus-1.5 model This repository contains the UI-Venus model from the report UI-Venus-1.5 Technical Report. UI-Venus 1.5 is a unified, end-to-end GUI Agent designed for robust real-world applications. The model family includes two dense (2B/8B) and one MoE (30B-A3B) variants to…

Multimodal 2.0B ↓ 11.8K
SmolLM-360M-Instruct logo
SmolLM-360M-Instruct
HuggingFaceTB

Model Summary Chat with the model at: https://huggingface.co/spaces/HuggingFaceTB/instant-smol

Open Source 0.36B ↓ 11.5K