AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

562 models for "Vision" Compare
InternVL3-2B-hf logo
InternVL3-2B-hf
OpenGVLab

InternVL3-2B Transformers 🤗 Implementation

Multimodal 2.09B ↓ 11.4K
InternVL3_5-30B-A3B logo
InternVL3_5-30B-A3B
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 30.85B ↓ 11K
SmolLM-1.7B-Instruct logo
SmolLM-1.7B-Instruct
HuggingFaceTB

SmolLM is a series of small language models available in three sizes: 135M, 360M, and 1.7B parameters.

Open Source 1.7B ↓ 10.8K
Molmo-72B-0924 logo
Molmo-72B-0924
allenai

Molmo is a family of open vision-language models developed by the Allen Institute for AI. Molmo models are trained on PixMo, a dataset of 1 million, highly-curated image-text pairs. It has state-of-the-art performance among multimodal models with a similar size while being fully…

Multimodal 72.0B ↓ 10.6K
InternVL3-14B-AWQ logo
InternVL3-14B-AWQ
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 15.12B ↓ 10.5K
UI-Venus-2-9B logo
UI-Venus-2-9B
inclusionAI

UI-Venus-2 is a general-purpose foundation GUI agent designed to operate across mobile applications, web platforms, and desktop operating systems through a unified closed-loop reasoning–action framework: the agent observes the current interface, reasons about the task state, exec…

Open Source 9.0B ↓ 10K
OLMoE-1B-7B-0125 logo
OLMoE-1B-7B-0125
allenai

OLMoE-1B-7B is a Mixture-of-Experts LLM with 1B active and 7B total parameters released in January 2025 (0125) that is 100% open-source. It is an improved version of OLMoE-09-24, see the paper appendix for details.

Open Source 6.92B ↓ 9.9K
LFM2.5-1.2B-JP-202606 logo
LFM2.5-1.2B-JP-202606
LiquidAI

LFM2.5-1.2B-JP-202606 is our latest general purpose Japanese chat model, delivering significant improvements in knowledge, instruction following, math, code, and tool-use over both the models of comparable size and LFM2.5-1.2B-JP. It sets a new benchmark for state-of-the-art perf…

Open Source 1.2B ↓ 9.9K
OLMo-2-0325-32B-Instruct logo
OLMo-2-0325-32B-Instruct
allenai

OLMo 2 32B Instruct March 2025 is post-trained variant of the OLMo-2 32B March 2025 model, which has undergone supervised finetuning on an OLMo-specific variant of the Tülu 3 dataset, further DPO training on this dataset, and final RLVR training on this dataset. Tülu 3 is designe…

Open Source 32.0B ↓ 9.9K
InternVL3_5-14B logo
InternVL3_5-14B
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 15.1B ↓ 9.3K
granite-3.1-3b-a800m-instruct logo
granite-3.1-3b-a800m-instruct
ibm-granite

Model Summary: Granite-3.1-3B-A800M-Instruct is a 3B parameter long-context instruct model finetuned from Granite-3.1-3B-A800M-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets tailored for solving lon…

Open Source 3.0B ↓ 9.1K
MolmoPoint-8B logo
MolmoPoint-8B
allenai

MolmoPoint-8B MolmoPoint-8B is a fully-open VLM developed by the Allen Institute for AI (Ai2) that support image, video and multi-image understanding and grounding. It has new pointing mechansim that improves image pointing, video pointing, and video tracking, see our technical r…

Multimodal 9.0B ↓ 8.7K