AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

562 models for "Vision" Compare
Qwen: Qwen3 VL 235B A22B Instruct logo
Qwen: Qwen3 VL 235B A22B Instruct
qwen

Qwen3-VL-235B-A22B Instruct is an open-weight multimodal model that unifies strong text generation with visual understanding across images and video. The Instruct model targets general vision-language use (VQA, document parsing, chart/table...

Multimodal 235.0B
OpenAI: GPT-4 Turbo (batch) logo
OpenAI: GPT-4 Turbo (batch)
openai

The latest GPT-4 Turbo model with vision capabilities. Vision requests can now use JSON mode and function calling. Training data: up to December 2023.

Closed Source
OpenAI: GPT-4 Turbo logo
OpenAI: GPT-4 Turbo
openai

The latest GPT-4 Turbo model with vision capabilities. Vision requests can now use JSON mode and function calling. Training data: up to December 2023.

Closed Source
Qwen: Qwen3 VL 32B Instruct logo
Qwen: Qwen3 VL 32B Instruct
qwen

Qwen3-VL-32B-Instruct is a large-scale multimodal vision-language model designed for high-precision understanding and reasoning across text, images, and video. With 32 billion parameters, it combines deep visual perception with advanced text...

Multimodal
Qwen: Qwen3.5-Flash logo
Qwen: Qwen3.5-Flash
qwen

The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. Compared to the...

Closed Source
Reka Edge logo
Reka Edge
rekaai

Reka Edge is an extremely efficient 7B multimodal vision-language model that accepts image/video+text inputs and generates text outputs. This model is optimized specifically to deliver industry-leading performance in image understanding,...

Closed Source
Z.ai: GLM 5V Turbo logo
Z.ai: GLM 5V Turbo
z-ai

GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image, video, and text inputs, excels at long-horizon planning, complex coding,...

Closed Source
Perceptron: Perceptron Mk1 logo
Perceptron: Perceptron Mk1
perceptron

Perceptron Mk1 (Mark One) is Perceptron's highest-quality vision-language model for video and embodied reasoning.** It accepts image and video inputs paired with natural language queries, and produces detailed visual understanding...

Closed Source
Qwen: Qwen3.7 Flash logo
Qwen: Qwen3.7 Flash
qwen

Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer interaction, with strengths in object recognition, spatial understanding, and real-world...

Closed Source
bilingual-gpt-neox-4b-minigpt4 logo
bilingual-gpt-neox-4b-minigpt4
rinna

Overview This repository provides an English-Japanese bilingual multimodal conversational model like MiniGPT-4 by combining GPT-NeoX model of 3.8 billion parameters and BLIP-2.

Open Source 4.0B