AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

562 models for "Vision" Compare
LocateAnything-3B logo
LocateAnything-3B
nvidia

LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding

Multimodal 3.0B ↓ 210.3K
Florence-2-base-ft logo
Florence-2-base-ft
microsoft

Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks

Multimodal 0.23B ↓ 193.3K
gemma-3n-E2B-it logo
gemma-3n-E2B-it
google

Multimodal 2.0B ↓ 170.8K
Cosmos-Reason2-8B logo
Cosmos-Reason2-8B
nvidia

Multimodal 8.0B ↓ 167K
GLM-4.1V-9B-Thinking logo
GLM-4.1V-9B-Thinking
zai-org

📖 View the GLM-4.1V-9B-Thinking paper . 📍 Using GLM-4.1V-9B-Thinking API at Zhipu Foundation Model Open Platform

Open Source 9.0B ↓ 164.2K
Phi-3-mini-128k-instruct logo
Phi-3-mini-128k-instruct
microsoft

🎉 Phi-4 : [multimodal-instruct onnx]; [mini-instruct onnx]

Open Source ↓ 160.7K
Kimi-K2.6-NVFP4 logo
Kimi-K2.6-NVFP4
nvidia

Description: The NVIDIA Kimi-K2.6-NVFP4 model is the quantized version of the Moonshot AI's Kimi-K2.6 model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Kimi-K2.6 NVFP4 model is qu…

Open Source ↓ 153.2K
NVIDIA-Nemotron-Parse-v1.1 logo
NVIDIA-Nemotron-Parse-v1.1
nvidia

NVIDIA Nemotron Parse v1.1 is designed to understand document semantics and extract text and tables elements with spatial grounding. Given an image, NVIDIA Nemotron Parse v1.1 produces structured annotations, including formatted text, bounding-boxes and the corresponding semantic…

Multimodal 0.885B ↓ 148.9K
SmolVLM2-2.2B-Instruct logo
SmolVLM2-2.2B-Instruct
HuggingFaceTB

SmolVLM2-2.2B is a lightweight multimodal model designed to analyze video content. The model processes videos, images, and text inputs to generate text outputs - whether answering questions about media files, comparing visual content, or transcribing text from images. Despite its…

Multimodal 2.2B ↓ 140.2K
Phi-3.5-MoE-instruct logo
Phi-3.5-MoE-instruct
microsoft

Phi-3.5-MoE is a lightweight, state-of-the-art open model built upon datasets used for Phi-3 - synthetic data and filtered publicly available documents - with a focus on very high-quality, reasoning dense data. The model supports multilingual and comes with 128K context length (i…

Open Source ↓ 138.8K
Nemotron-Labs-Diffusion-8B logo
Nemotron-Labs-Diffusion-8B
nvidia

Nemotron-Labs-Diffusion is a tri-mode language model that supports both AR decoding and diffusion-based parallel decoding by simply switching the attention pattern of the same model during inference. The synergy between these two modes enables a third mode, called self-speculatio…

Open Source 8.0B ↓ 136.5K
SmolVLM-500M-Instruct logo
SmolVLM-500M-Instruct
HuggingFaceTB

SmolVLM-500M is a tiny multimodal model, member of the SmolVLM family. It accepts arbitrary sequences of image and text inputs to produce text outputs. It's designed for efficiency. SmolVLM can answer questions about images, describe visual content, or transcribe text. Its lightw…

Multimodal 0.5B ↓ 133K