AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

562 models for "Vision" Compare
Nemotron-Labs-Diffusion-8B-Base logo
Nemotron-Labs-Diffusion-8B-Base
nvidia

Nemotron-Labs-Diffusion is a tri-mode language model that supports both AR decoding and diffusion-based parallel decoding by simply switching the attention pattern of the same model during inference. The synergy between these two modes enables a third mode, called self-speculatio…

Open Source 8.0B ↓ 70.7K
Nemotron-Labs-Diffusion-3B logo
Nemotron-Labs-Diffusion-3B
nvidia

Nemotron-Labs-Diffusion is a tri-mode language model that supports both AR decoding and diffusion-based parallel decoding by simply switching the attention pattern of the same model during inference. The synergy between these two modes enables a third mode, called self-speculatio…

Open Source 3.0B ↓ 70.7K
LFM2-VL-1.6B logo
LFM2-VL-1.6B
LiquidAI

LFM2‑VL is Liquid AI's first series of multimodal models, designed to process text and images with variable resolutions. Built on the LFM2 backbone, it is optimized for low-latency and edge AI applications.

Multimodal 1.6B ↓ 66.2K
Llama-Guard-4-12B logo
Llama-Guard-4-12B
meta-llama

Open Source 12.0B ↓ 61.1K
OLMo-2-0425-1B-Instruct logo
OLMo-2-0425-1B-Instruct
allenai

OLMo 2 1B Instruct April 2025 is post-trained variant of the allenai/OLMo-2-0425-1B-RLVR1 model, which has undergone supervised finetuning on an OLMo-specific variant of the Tülu 3 dataset, further DPO training on this dataset, and final RLVR training on this dataset. Tülu 3 is d…

Open Source 1.0B ↓ 60.8K
granite-3.1-8b-instruct logo
granite-3.1-8b-instruct
ibm-granite

Model Summary: Granite-3.1-8B-Instruct is a 8B parameter long-context instruct model finetuned from Granite-3.1-8B-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets tailored for solving long context pr…

Open Source 8.0B ↓ 55.3K
udop-large logo
udop-large
microsoft

The UDOP model was proposed in Unifying Vision, Text, and Layout for Universal Document Processing by Zineng Tang, Ziyi Yang, Guoxin Wang, Yuwei Fang, Yang Liu, Chenguang Zhu, Michael Zeng, Cha Zhang, Mohit Bansal.

Multimodal 0.794B ↓ 55K
granite-4.0-tiny-preview logo
granite-4.0-tiny-preview
ibm-granite

Model Summary: Granite-4-Tiny-Preview is a 7B parameter fine-grained hybrid mixture-of-experts (MoE) instruct model fine-tuned from Granite-4.0-Tiny-Base-Preview using a combination of open source instruction datasets with permissive license and internally collected synthetic dat…

Open Source 7.0B ↓ 54.1K
granite-3.2-8b-instruct logo
granite-3.2-8b-instruct
ibm-granite

Model Summary: Granite-3.2-8B-Instruct is an 8-billion-parameter, long-context AI model fine-tuned for thinking capabilities. Built on top of Granite-3.1-8B-Instruct, it has been trained using a mix of permissively licensed open-source datasets and internally generated synthetic…

Open Source 8.17B ↓ 52.9K
step3 logo
step3
stepfun-ai

📰   Step3 Model Blog         📄   Step3 System Blog

Open Source ↓ 52.1K
codegen2-16B_P logo
codegen2-16B_P
Salesforce

CodeGen2 is a family of autoregressive language models for program synthesis , introduced in the paper:

Code 16.0B ↓ 49.9K
blip2-flan-t5-xl logo
blip2-flan-t5-xl
Salesforce

BLIP-2 model, leveraging Flan T5-xl (a large language model). It was introduced in the paper BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models by Li et al. and first released in this repository.

Multimodal 4.0B ↓ 49.9K