AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

562 models for "Vision" Compare
Baichuan2-13B-Chat logo
Baichuan2-13B-Chat
baichuan-inc

🦉GitHub 💬WeChat 百川API支持搜索增强和192K长窗口,新增百川搜索增强知识库、限时免费! 🚀 百川大模型在线对话平台 已正式向公众开放 🎉

Open Source 13.0B ↓ 8.7K
d1-omni-600M logo
d1-omni-600M
LiquidAI

d1-omni-600M is a 600M parameter decision model built on LFM2.5-Encoder-350M. You give it a state (text or JSON, with images or a voice clip) and a set of named questions. It returns typed answers with zero output tokens : every answer is read directly from the model's distributi…

Multimodal 0.587B ↓ 8.5K
granite-3.0-2b-instruct logo
granite-3.0-2b-instruct
ibm-granite

Model Summary: Granite-3.0-2B-Instruct is a 2B parameter model finetuned from Granite-3.0-2B-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets. This model is developed using a diverse set of techniques…

Open Source 2.5B ↓ 8.4K
OLMo-2-0425-1B-SFT logo
OLMo-2-0425-1B-SFT
allenai

OLMo 2 1B SFT April 2025 is post-trained variant of the allenai/OLMo-2-0425-1B model, which has undergone supervised finetuning on an OLMo-specific variant of the Tülu 3 dataset. Tülu 3 is designed for state-of-the-art performance on a diversity of tasks in addition to chat, such…

Open Source 1.0B ↓ 8.4K
GLM-4.6V-FP8 logo
GLM-4.6V-FP8
zai-org

This model is part of the GLM-V family of models, introduced in the paper GLM-4.1V-Thinking and GLM-4.5V: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

Open Source ↓ 8.3K
InternVL-Chat-V1-5-AWQ logo
InternVL-Chat-V1-5-AWQ
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 Mini-InternVL\]](https://arxiv.org/abs/2410.16261) [\[📜 InternVL 2.5\]](https://huggingface.co/…

Multimodal 25.5B ↓ 8.3K
InternVL2_5-4B logo
InternVL2_5-4B
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 Mini-InternVL\]](https://arxiv.org/abs/2410.16261) [\[📜 InternVL 2.5\]](https://huggingface.co/…

Multimodal 3.7B ↓ 8.3K
granite-3.1-2b-instruct logo
granite-3.1-2b-instruct
ibm-granite

Model Summary: Granite-3.1-2B-Instruct is a 2B parameter long-context instruct model finetuned from Granite-3.1-2B-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets tailored for solving long context pr…

Open Source 2.0B ↓ 8.1K
InternVL3-78B logo
InternVL3-78B
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 78.4B ↓ 8.1K
OLMo-2-0425-1B-DPO logo
OLMo-2-0425-1B-DPO
allenai

OLMo 2 1B DPO April 2025 is post-trained variant of the allenai/OLMo-2-0425-1B-SFT model, which has undergone supervised finetuning on an OLMo-specific variant of the Tülu 3 dataset and further DPO training on this dataset. Tülu 3 is designed for state-of-the-art performance on a…

Open Source 1.0B ↓ 7.8K
Mini-InternVL2-4B-DA-Medical logo
Mini-InternVL2-4B-DA-Medical
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[🆕 Blog\]](https://internvl.github.io/blog/) [\[📜 Mini-InternVL\]](https://arxiv.org/abs/2410.16261) [\[📜 InternVL 1.0\]](https://arxiv.org/abs/2312.14238) [\[📜 InternVL 1.5\]](https://arxiv.org/abs/2404.16821) [\[📜 InternVL…

Multimodal 4.0B ↓ 7.8K
internlm-xcomposer2-7b logo
internlm-xcomposer2-7b
internlm

InternLM-XComposer2 is a vision-language large model (VLLM) based on InternLM2 for advanced text-image comprehension and composition.

Open Source 7.0B ↓ 7.6K