AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

562 models for "Vision" Compare
MiniCPM-V-4_5 logo
MiniCPM-V-4_5
openbmb

A GPT-4o Level MLLM for Single Image, Multi Image and High-FPS Video Understanding on Your Phone

Multimodal 8.7B ↓ 93K
paligemma-3b-mix-224 logo
paligemma-3b-mix-224
google

Multimodal 2.92B ↓ 91.8K
deepseek-vl2-tiny logo
deepseek-vl2-tiny
deepseek-ai

Introducing DeepSeek-VL2, an advanced series of large Mixture-of-Experts (MoE) Vision-Language Models that significantly improves upon its predecessor, DeepSeek-VL. DeepSeek-VL2 demonstrates superior capabilities across various tasks, including but not limited to visual question…

Multimodal 3.37B ↓ 87.4K
Qianfan-OCR logo
Qianfan-OCR
baidu

A Unified End-to-End Model for Document Intelligence

Multimodal 4.0B ↓ 86.6K
granite-3.0-1b-a400m-instruct logo
granite-3.0-1b-a400m-instruct
ibm-granite

Model Summary: Granite-3.0-1B-A400M-Instruct is an 1B parameter model finetuned from Granite-3.0-1B-A400M-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets. This model is developed using a diverse set…

Open Source 1.3B ↓ 86.6K
Llama-3.1-Nemotron-Nano-VL-8B-V1 logo
Llama-3.1-Nemotron-Nano-VL-8B-V1
nvidia

Llama Nemotron Nano VL is a leading document intelligence vision language model (VLMs) that enables the ability to query and summarize images from the physical or virtual world. Llama Nemotron Nano VL is deployable in the data center, cloud and at the edge, including Jetson Orin…

Multimodal 9.0B ↓ 82.6K
granite-3.3-8b-instruct logo
granite-3.3-8b-instruct
ibm-granite

Model Summary: Granite-3.3-8B-Instruct is a 8-billion parameter 128K context length language model fine-tuned for improved reasoning and instruction-following capabilities. Built on top of Granite-3.3-8B-Base, the model delivers significant gains on benchmarks for measuring gener…

Open Source 8.0B ↓ 81.8K
InternVL2_5-4B-MPO logo
InternVL2_5-4B-MPO
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 4.0B ↓ 80.1K
InternVL3-8B-hf logo
InternVL3-8B-hf
OpenGVLab

InternVL3-8B Transformers 🤗 Implementation

Multimodal 8.0B ↓ 74.6K
InternVL2_5-2B logo
InternVL2_5-2B
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 Mini-InternVL\]](https://arxiv.org/abs/2410.16261) [\[📜 InternVL 2.5\]](https://huggingface.co/…

Multimodal 2.2B ↓ 74.1K
Hunyuan-A13B-Instruct logo
Hunyuan-A13B-Instruct
tencent

🤗  Hugging Face       🖥️  Official Website       🕖  HunyuanAPI       🕹️  Demo       🤖  ModelScope

Open Source 13.0B ↓ 72K
ctrl logo
ctrl
Salesforce

1. Model Details 2. Uses 3. Bias, Risks, and Limitations 4. Training 5. Evaluation 6. Environmental Impact 7. Technical Specifications 8. Citation 9. Model Card Authors 10. How To Get Started With the Model

Open Source 1.6B ↓ 70.9K