AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

562 models for "Vision" Compare
OLMo-7B logo
OLMo-7B
allenai

For transformers versions v4.40.0 or newer, we suggest using OLMo 7B HF instead.

Open Source 7.0B ★ 3.0 ↓ 4.9K
granite-4.0-1b-base logo
granite-4.0-1b-base
ibm-granite

Model Summary: Granite-4.0-1B-Base is a lightweight decoder-only language model designed for scenarios where efficiency and speed are critical. They can run on resource-constrained devices such as smartphones or IoT hardware, enabling offline and privacy-preserving applications.…

Open Source 1.6B ★ 2.0 ↓ 44.7K
granite-4.0-1b logo
granite-4.0-1b
ibm-granite

Model Summary: Granite-4.0-1B is a lightweight instruct model finetuned from Granite-4.0-1B-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets. This model is developed using a diverse set of techniques…

Open Source 1.0B ★ 2.0 ↓ 11.7K
granite-4.0-h-micro logo
granite-4.0-h-micro
ibm-granite

📣 Update [10-07-2025]: Added a default system prompt to the chat template to guide the model towards more professional, accurate, and safe responses.

Open Source 3.0B ★ 2.0 ↓ 7.9K
LFM2.5-VL-1.6B-Extract logo
LFM2.5-VL-1.6B-Extract
LiquidAI

LFM2.5-VL-1.6B-Extract extracts user-defined fields from images and returns them as JSON . It is Liquid AI's first vision model in the Liquid Nanos collection—compact, task-specific models built for production workflows—and extends the Extract family alongside LFM2-1.2B-Extract f…

Multimodal 1.6B ★ 1.0 ↓ 1.9K
Qwen3-VL-8B-Instruct logo
Qwen3-VL-8B-Instruct
Qwen

Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.

Multimodal 8.0B ↓ 9.7M
Qwen2.5-VL-7B-Instruct logo
Qwen2.5-VL-7B-Instruct
Qwen

--- license: apache-2.0 language: - en pipeline tag: image-text-to-text tags: - multimodal library name: transformers ---

Multimodal 7.0B ↓ 5.4M
Qwen3-VL-4B-Instruct logo
Qwen3-VL-4B-Instruct
Qwen

Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.

Multimodal 4.0B ↓ 3.2M
gemma-3-1b-it logo
gemma-3-1b-it
google

Open Source 1.0B ↓ 3.1M
Florence-2-base logo
Florence-2-base
microsoft

Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks

Multimodal 0.23B ↓ 3.1M
Qwen3-VL-2B-Instruct logo
Qwen3-VL-2B-Instruct
Qwen

Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.

Multimodal 2.0B ↓ 2.7M
Qwen2.5-VL-3B-Instruct logo
Qwen2.5-VL-3B-Instruct
Qwen

--- license name: qwen-research license link: https://huggingface.co/Qwen/Qwen2.5-VL-3B-Instruct/blob/main/LICENSE language: - en pipeline tag: image-text-to-text tags: - multimodal library name: transformers ---

Multimodal 3.0B ↓ 2.3M