AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

562 models for "Vision" Compare
DeepSeek: DeepSeek V4 Flash Vision Exp logo
DeepSeek: DeepSeek V4 Flash Vision Exp
deepseek

DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of [DeepSeek V4 Flash 0731](https://openrouter.ai/deepseek/deepseek-v4-flash-0731) from DeepSeek, adding image understanding while matching the base model on text capabilities including agents,...

Multimodal ★ 51.0
DeepSeek-V4-Flash-Vision-Exp logo
DeepSeek-V4-Flash-Vision-Exp
deepseek-ai

We are excited to introduce DeepSeek-V4-Flash-Vision-Exp , our first experimental multimodal model in the DeepSeek-V4 family. It builds on the DeepSeek-V4-Flash architecture by incorporating visual modules and undergoing continued training to unlock visual understanding capabilit…

Multimodal ★ 35.0 ↓ 754.2K
command-a-vision-07-2025 logo
command-a-vision-07-2025
CohereLabs

Multimodal 112.0B ★ 13.0 ↓ 19K
Llama-3.2-90B-Vision-Instruct logo
Llama-3.2-90B-Vision-Instruct
meta-llama

Multimodal 90.0B ★ 6.0 ↓ 157.9K
Llama-3.2-90B-Vision logo
Llama-3.2-90B-Vision
meta-llama

Multimodal 90.0B ★ 6.0 ↓ 48
Llama-3.2-11B-Vision-Instruct logo
Llama-3.2-11B-Vision-Instruct
meta-llama

Multimodal 11.0B ★ 3.0 ↓ 101.8K
Llama-3.2-11B-Vision logo
Llama-3.2-11B-Vision
meta-llama

Multimodal 11.0B ★ 3.0 ↓ 10K
Phi-3.5-vision-instruct logo
Phi-3.5-vision-instruct
microsoft

Phi-3.5-vision is a lightweight, state-of-the-art open multimodal model built upon datasets which include - synthetic data and filtered publicly available websites - with a focus on very high-quality, reasoning dense data both on text and vision. The model belongs to the Phi-3 mo…

Multimodal 4.2B ↓ 740.1K
granite-vision-4.1-4b logo
granite-vision-4.1-4b
ibm-granite

Model Summary: Granite Vision 4.1 4B is a vision-language model (VLM) that delivers frontier-level performance on structured document extraction tasks — chart extraction, table extraction, and semantic key-value pair extraction — in a compact 4B parameter footprint, providing a l…

Multimodal 4.0B ↓ 179.4K
North-Micro-Vision-Instruct logo
North-Micro-Vision-Instruct
CohereLabs

North Micro Vision Instruct is a 2.4B-parameter open-weight vision-language model with native-resolution image support, released under the Apache 2.0 license. It is designed as a compact foundation for prototyping, task-specific fine-tuning, and specialized multimodal application…

Multimodal 2.4B ↓ 134K
granite-4.0-3b-vision logo
granite-4.0-3b-vision
ibm-granite

Model Summary: Granite-4.0-3B-Vision is a vision-language model (VLM) designed for enterprise-grade document data extraction. It focuses on specialized, complex extraction tasks that ultracompact models often struggle with:

Multimodal 4.0B ↓ 56.3K
Phi-3-vision-128k-instruct logo
Phi-3-vision-128k-instruct
microsoft

🎉 Phi-3.5 : [[mini-instruct]](https://huggingface.co/microsoft/Phi-3.5-mini-instruct); [[MoE-instruct]](https://huggingface.co/microsoft/Phi-3.5-MoE-instruct) ; [[vision-instruct]](https://huggingface.co/microsoft/Phi-3.5-vision-instruct)

Multimodal ↓ 52K