AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

562 models for "Vision" Compare
InternVL3_5-8B-Pretrained logo
InternVL3_5-8B-Pretrained
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 8.0B ↓ 1.2K
Qwen-SEA-LION-v4-4B-VL logo
Qwen-SEA-LION-v4-4B-VL
aisingapore

SEA-LION is a collection of Large Language Models (LLMs) which have been pretrained and instruct-tuned for the Southeast Asia (SEA) region.

Multimodal 4.0B ↓ 1.1K
Intern-S2-397B logo
Intern-S2-397B
internlm

💻Github Repo • 🤗HF Model Collections • 🤖ModelScope Collections • 💬Online Chat

Multimodal 397.0B ↓ 1.1K
Ring-mini-linear-2.0 logo
Ring-mini-linear-2.0
inclusionAI

📖 Technical Report &nbsp&nbsp &nbsp&nbsp 🤗 Hugging Face &nbsp&nbsp &nbsp&nbsp🤖 ModelScope

Open Source 16.4B ↓ 1.1K
AutoGLM-Phone-9B-Multilingual logo
AutoGLM-Phone-9B-Multilingual
zai-org

⚠️ This project is intended for research and educational purposes only . Any use for illegal data access, system interference, or unlawful activities is strictly prohibited. Please review our Terms of Use carefully.

Open Source 9.0B ↓ 1.1K
InternVL3_5-2B-Instruct logo
InternVL3_5-2B-Instruct
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 2.35B ↓ 1.1K
VisualPRM-8B logo
VisualPRM-8B
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📂 Evaluation Code\]](https://github.com/open-compass/VLMEvalKit/pull/854/files) [\[📜 Paper\]](https://arxiv.org/abs/2503.10291) [\[🆕 Blog\]](https://internvl.github.io/blog/2025-03-13-VisualPRM/) [\[🤗 model\]](https://huggi…

Multimodal 8.0B ↓ 1.1K
Qwen-SEA-LION-v4.5-27B-IT logo
Qwen-SEA-LION-v4.5-27B-IT
aisingapore

SEA-LION is a collection of Large Language Models (LLMs) which have been pretrained and instruct-tuned for the Southeast Asia (SEA) region.

Multimodal 27.0B ↓ 1.1K
InternVL3_5-241B-A28B-Instruct logo
InternVL3_5-241B-A28B-Instruct
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 241.0B ↓ 1K
InternVL3_5-1B-Flash logo
InternVL3_5-1B-Flash
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 1.1B ↓ 1K
cogvlm-chat-hf logo
cogvlm-chat-hf
zai-org

CogVLM 是一个强大的开源视觉语言模型(VLM)。CogVLM-17B 拥有 100 亿视觉参数和 70 亿语言参数,在 10 个经典跨模态基准测试上取得了 SOTA 性能,包括 NoCaps、Flicker30k captioning、RefCOCO、RefCOCO+、RefCOCOg、Visual7W、GQA、ScienceQA、VizWiz VQA 和 TDIUC,而在 VQAv2、OKVQA、TextVQA、COCO captioning 等方面则排名第二,超越或与 PaLI-X 55B 持平。您可以通过线上 demo 体验 CogVLM 多…

Multimodal ↓ 984
blip2-flan-t5-xxl logo
blip2-flan-t5-xxl
Salesforce

BLIP-2 model, leveraging Flan T5-xxl (a large language model). It was introduced in the paper BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models by Li et al. and first released in this repository.

Open Source ↓ 977