AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

562 models for "Vision" Compare
granite-vision-3.2-2b logo
granite-vision-3.2-2b
ibm-granite

Model Summary: granite-vision-3.2-2b is a compact and efficient vision-language model, specifically designed for visual document understanding, enabling automated content extraction from tables, charts, infographics, plots, diagrams, and more. The model was trained on a meticulou…

Multimodal 2.0B ↓ 4.3K
aya-vision-8b logo
aya-vision-8b
CohereLabs

Multimodal 8.0B ↓ 2.8K
VisionReward-Video logo
VisionReward-Video
zai-org

Introduction We present VisionReward, a general strategy to aligning visual generation models——both image and video generation——with human preferences through a fine-grainedand multi-dimensional framework. We decompose human preferences in images and videos into multiple dimensio…

Multimodal 13.0B ↓ 2.5K
Llama-Guard-3-11B-Vision logo
Llama-Guard-3-11B-Vision
meta-llama

Multimodal 11.0B ↓ 2.1K
Nous-Hermes-2-Vision-Alpha logo
Nous-Hermes-2-Vision-Alpha
NousResearch

In the tapestry of Greek mythology, Hermes reigns as the eloquent Messenger of the Gods, a deity who deftly bridges the realms through the art of communication. It is in homage to this divine mediator that I name this advanced LLM "Hermes," a system crafted to navigate the comple…

Multimodal ↓ 373
aya-vision-32b logo
aya-vision-32b
CohereLabs

Multimodal 32.0B ↓ 249
Anthropic: Claude Sonnet 5.5 logo
Anthropic: Claude Sonnet 5.5
anthropic

Claude Sonnet 5.5 is Anthropic's Sonnet-class model for well-scoped everyday work, succeeding Claude Sonnet 5 as a direct upgrade. It is especially strong at building features, fixing bugs, and producing...

Closed Source ★ 56.0
OpenAI: GPT Astra Latest logo
OpenAI: GPT Astra Latest
~openai

This model always redirects to the latest model in the GPT Astra family.

Closed Source ★ 53.0
DeepSeek: DeepSeek V4 Flash 0731 (batch) logo
DeepSeek: DeepSeek V4 Flash 0731 (batch)
deepseek

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....

Closed Source ★ 52.0
MiMo-V2.6-Pro-RL logo
MiMo-V2.6-Pro-RL
XiaomiMiMo

🤗 HuggingFace   📰 Blog   🎨 Xiaomi MiMo API Platform   🗨️ Xiaomi MiMo Studio   💻 Xiaomi MiMo Desktop  

Open Source ★ 46.0 ↓ 93.5K
MiMo-V2.6-Pro-MOPD logo
MiMo-V2.6-Pro-MOPD
XiaomiMiMo

🤗 HuggingFace   📰 Blog   🎨 Xiaomi MiMo API Platform   🗨Xiaomi MiMo Studio   💻 Xiaomi MiMo Desktop  

Open Source ★ 46.0 ↓ 3.6K
🤖
Xiaomi: MiMo-V2.6-Pro
xiaomi

MiMo-V2.6-Pro is the flagship foundation model developed by Xiaomi. Built at a scale of over 1T parameters, it is designed to push the ceiling of capability for the most demanding...

Open Source 1020.0B ★ 46.0