AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

562 models for "Vision" Compare
Qwen3.5-27B logo
Qwen3.5-27B
Qwen

[!Note] This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.

Open Source 27.0B ↓ 1.9M
Qwen2.5-VL-32B-Instruct-AWQ logo
Qwen2.5-VL-32B-Instruct-AWQ
Qwen

Latest Updates: In addition to the original formula, we have further enhanced Qwen2.5-VL-32B's mathematical and problem-solving abilities through reinforcement learning. This has also significantly improved the model's subjective user experience, with response styles adjusted to…

Multimodal 32.0B ↓ 1.8M
Qwen3-VL-8B-Instruct-FP8 logo
Qwen3-VL-8B-Instruct-FP8
Qwen

This repository contains an FP8 quantized version of the Qwen3-VL-8B-Instruct model. The quantization method is fine-grained fp8 quantization with block size of 128, and its performance metrics are nearly identical to those of the original BF16 model. Enjoy!

Multimodal 8.0B ↓ 1.8M
Qwen2.5-VL-7B-Instruct-AWQ logo
Qwen2.5-VL-7B-Instruct-AWQ
Qwen

In the past five months since Qwen2-VL’s release, numerous developers have built new models on the Qwen2-VL vision-language models, providing us with valuable feedback. During this period, we focused on building more useful vision-language models. Today, we are excited to introdu…

Multimodal 7.0B ↓ 1.7M
Qwen2-VL-7B-Instruct-AWQ logo
Qwen2-VL-7B-Instruct-AWQ
Qwen

We're excited to unveil Qwen2-VL , the latest iteration of our Qwen-VL model, representing nearly a year of innovation.

Multimodal 7.0B ↓ 1.6M
Unlimited-OCR logo
Unlimited-OCR
baidu

Welcome the Era of One-shot Long-horizon Parsing.

Multimodal 3.336B ↓ 1.3M
Qwen2-VL-2B-Instruct logo
Qwen2-VL-2B-Instruct
Qwen

We're excited to unveil Qwen2-VL , the latest iteration of our Qwen-VL model, representing nearly a year of innovation.

Multimodal 2.0B ↓ 1.3M
SmolVLM2-500M-Video-Instruct logo
SmolVLM2-500M-Video-Instruct
HuggingFaceTB

SmolVLM2-500M-Video is a lightweight multimodal model designed to analyze video content. The model processes videos, images, and text inputs to generate text outputs - whether answering questions about media files, comparing visual content, or transcribing text from images. Despi…

Multimodal 0.5B ↓ 1.3M
gemma-3-4b-it logo
gemma-3-4b-it
google

Multimodal 4.0B ↓ 1.2M
HunyuanOCR logo
HunyuanOCR
tencent

HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better

Multimodal 1.0B ↓ 939.5K
medgemma-4b-it logo
medgemma-4b-it
google

Multimodal 4.0B ↓ 849K
Qwen3-VL-235B-A22B-Instruct logo
Qwen3-VL-235B-A22B-Instruct
Qwen

Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.

Multimodal 235.0B ↓ 825.5K