AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

562 models for "Vision" Compare
NVIDIA-Nemotron-Nano-12B-v2-VL-NVFP4-QAD logo
NVIDIA-Nemotron-Nano-12B-v2-VL-NVFP4-QAD
nvidia

NVIDIA-Nemotron-Nano-VL-12B-V2-FP4-QAD is the quantized version of the NVIDIA Nemotron Nano VL V2 model, which is an auto-regressive vision language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Nemotron Nano VL FP4 QAD…

Multimodal 12.6B ★ 7.0 ↓ 169.2K
Olmo-3.1-32B-Think logo
Olmo-3.1-32B-Think
allenai

We introduce Olmo 3, a new family of 7B and 32B models both Instruct and Think variants. Long chain-of-thought thinking improves reasoning tasks like math and coding.

Open Source 32.0B ★ 7.0 ↓ 19.2K
Qwen3.5-0.8B logo
Qwen3.5-0.8B
Qwen

[!Note] This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc. In light of its parameter scale, the intende…

Open Source 0.8B ★ 6.0 ↓ 2.5M
Olmo-3-7B-Think logo
Olmo-3-7B-Think
allenai

We introduce Olmo 3, a new family of 7B and 32B models both Instruct and Think variants. Long chain-of-thought thinking improves reasoning tasks like math and coding.

Open Source 7.0B ★ 6.0 ↓ 489.1K
MiniCPM-V-4.6 logo
MiniCPM-V-4.6
openbmb

A Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone

Multimodal 1.3B ★ 6.0 ↓ 238.6K
granite-4.1-8b logo
granite-4.1-8b
ibm-granite

Model Summary: Granite-4.1-8B is a 8B parameter long-context instruct model finetuned from Granite-4.1-8B-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets. Granite 4.1 models have gone through an impr…

Open Source 8.0B ★ 6.0 ↓ 157.9K
MiniCPM-V-4.6-Thinking-GPTQ logo
MiniCPM-V-4.6-Thinking-GPTQ
openbmb

This repository hosts the GPTQ (W4A16, GPTQModel) quantized version of MiniCPM-V 4.6 Thinking. For the original BF16 weights and the full model card, please refer to openbmb/MiniCPM-V-4.6-Thinking.

Multimodal 1.3B ★ 6.0 ↓ 50.3K
MiniCPM-V-4 logo
MiniCPM-V-4
openbmb

A GPT-4V Level MLLM for Single Image, Multi Image and Video on Your Phone

Multimodal 4.1B ★ 6.0 ↓ 46.7K
MiniCPM-V-4.6-BNB logo
MiniCPM-V-4.6-BNB
openbmb

This repository hosts the bitsandbytes (NF4, 4-bit) quantized version of MiniCPM-V 4.6. For the original BF16 weights and the full model card, please refer to openbmb/MiniCPM-V-4.6.

Multimodal 1.3B ★ 6.0 ↓ 40.3K
MiniCPM-V-4.6-GPTQ logo
MiniCPM-V-4.6-GPTQ
openbmb

This repository hosts the GPTQ (W4A16, GPTQModel) quantized version of MiniCPM-V 4.6. For the original BF16 weights and the full model card, please refer to openbmb/MiniCPM-V-4.6.

Multimodal 1.3B ★ 6.0 ↓ 39.4K
Molmo-7B-D-0924 logo
Molmo-7B-D-0924
allenai

Molmo is a family of open vision-language models developed by the Allen Institute for AI. Molmo models are trained on PixMo, a dataset of 1 million, highly-curated image-text pairs. It has state-of-the-art performance among multimodal models with a similar size while being fully…

Multimodal 8.0B ★ 6.0 ↓ 25.1K
granite-4.0-h-small logo
granite-4.0-h-small
ibm-granite

📣 Update [10-07-2025]: Added a default system prompt to the chat template to guide the model towards more professional, accurate, and safe responses.

Open Source 32.0B ★ 6.0 ↓ 24.1K