AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

562 models for "Vision" Compare
deepseek-vl2 logo
deepseek-vl2
deepseek-ai

Introducing DeepSeek-VL2, an advanced series of large Mixture-of-Experts (MoE) Vision-Language Models that significantly improves upon its predecessor, DeepSeek-VL. DeepSeek-VL2 demonstrates superior capabilities across various tasks, including but not limited to visual question…

Multimodal 27.0B ↓ 7.5K
granite-3.0-8b-base logo
granite-3.0-8b-base
ibm-granite

Model Summary: Granite-3.0-8B-Base is a decoder-only language model to support a variety of text-to-text generation tasks. It is trained from scratch following a two-stage training strategy. In the first stage, it is trained on 10 trillion tokens sourced from diverse domains. Dur…

Open Source 8.0B ↓ 7.4K
internlm-xcomposer-7b logo
internlm-xcomposer-7b
internlm

InternLM-XComposer is a vision-language large model (VLLM) based on InternLM for advanced text-image comprehension and composition. InternLM-XComposer has serveal appealing properties:

Multimodal 7.0B ↓ 7.3K
InternVL3_5-30B-A3B-Instruct logo
InternVL3_5-30B-A3B-Instruct
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 30.85B ↓ 7.3K
InternVL3_5-2B logo
InternVL3_5-2B
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 2.3B ↓ 7.2K
granite-3.3-8b-base logo
granite-3.3-8b-base
ibm-granite

Granite-3.3-8B-Base is a decoder-only language model with a 128K token context window. It improves upon Granite-3.1-8B-Base by adding support for Fill-in-the-Middle (FIM) using specialized tokens, enabling the model to generate content conditioned on both prefix and suffix. This…

Open Source 8.1B ↓ 7.1K
InternVL2_5-1B-MPO logo
InternVL2_5-1B-MPO
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 0.9B ↓ 7K
santacoder logo
santacoder
bigcode

Play with the model on the SantaCoder Space Demo.

Code 1.1B ↓ 6.8K
Kimi-VL-A3B-Thinking-2506 logo
Kimi-VL-A3B-Thinking-2506
moonshotai

[!Note] This is an improved version of Kimi-VL-A3B-Thinking. Please consider using this updated model instead of the previous version.

Multimodal 3.0B ↓ 6.6K
InternVL3_5-38B logo
InternVL3_5-38B
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 38.4B ↓ 6.5K
solar-pro-preview-instruct logo
solar-pro-preview-instruct
upstage

Solar Pro Preview: The most intelligent LLM on a single GPU

Open Source ↓ 6.5K
InternVL2_5-1B logo
InternVL2_5-1B
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 Mini-InternVL\]](https://arxiv.org/abs/2410.16261) [\[📜 InternVL 2.5\]](https://huggingface.co/…

Multimodal 0.9B ↓ 6.4K