AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

562 models for "Vision" Compare
opencole-typographylmm-llava-v1.5-7b-lora logo
opencole-typographylmm-llava-v1.5-7b-lora
cyberagent

This model is based on LLaVA1.5-7b. The model is finetuned with LoRA on OpenCOLE1.0 dataset to generate text layouts.

Open Source 7.0B ↓ 65
Qwen-SEA-LION-v4-32B-IT-8BIT logo
Qwen-SEA-LION-v4-32B-IT-8BIT
aisingapore

Qwen-SEA-LION-v4-32B-IT-8BIT (GPTQ model)

Open Source 32.0B ↓ 65
Gemma-SEA-LION-v4-27B-IT-FP8-Dynamic logo
Gemma-SEA-LION-v4-27B-IT-FP8-Dynamic
aisingapore

Model Card for Gemma-SEA-LION-v4-27B-IT-FP8-Dynamic

Open Source 27.0B ↓ 59
SmolVLM-Instruct-DPO logo
SmolVLM-Instruct-DPO
HuggingFaceTB

SmolVLM is a compact open multimodal model that accepts arbitrary sequences of image and text inputs to produce text outputs. Designed for efficiency, SmolVLM can answer questions about images, describe visual content, create stories grounded on multiple images, or function as a…

Multimodal ↓ 39
Meta: Llama Guard 4 12B logo
Meta: Llama Guard 4 12B
meta-llama

Llama Guard 4 is a Llama 4 Scout-derived multimodal pretrained model, fine-tuned for content safety classification. Similar to previous versions, it can be used to classify content in both LLM...

Multimodal 12.0B
Mistral: Mistral Medium 3.1 (batch) logo
Mistral: Mistral Medium 3.1 (batch)
mistralai

Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost. It balances...

Closed Source
DeepSeek: DeepSeek V4 Flash Latest logo
DeepSeek: DeepSeek V4 Flash Latest
~deepseek

This model always redirects to the latest model in the DeepSeek V4 Flash family.

Open Source 284.0B
Nex AGI: Nex-N2.5-Pro logo
Nex AGI: Nex-N2.5-Pro
nex-agi

Nex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual feedback loop: it can explore codebases, implement multi-file...

Multimodal 397.0B
Nex AGI: Nex-N2.5-Mini logo
Nex AGI: Nex-N2.5-Mini
nex-agi

Nex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual feedback loop: it can explore codebases, implement multi-file...

Multimodal 35.0B
🤖
PrismML: Ternary Bonsai 2 27B
prism-ml

Bonsai 2 27B is a 27B-parameter reasoning model from PrismML derived from Qwen3.8-27B. It supports coding, mathematics, tool calling, and image understanding with a 262K-token context window. Ternary compression shrinks...

Open Source 27.36B
Qwen: Qwen3.8 Omni Flash logo
Qwen: Qwen3.8 Omni Flash
qwen

Qwen3.8 Omni Flash is an omni-modal reasoning model from Alibaba, the first Qwen model built around agentic capabilities with native audio-video understanding. It is suited for audio-video analysis and summarization,...

Closed Source 125.0B
Mistral: Mistral Large 4 logo
Mistral: Mistral Large 4
mistralai

Mistral Large 4 is a frontier multimodal (text and image input) model from Mistral AI built for reasoning, coding, and agentic workloads. It offers a 1M-token context window with up...

Multimodal 1050.0B