AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

562 models for "Vision" Compare
gemma-3-4b-pt logo
gemma-3-4b-pt
google

Open Source 4.0B ↓ 129.3K
olmOCR-2-7B-1025 logo
olmOCR-2-7B-1025
allenai

Full BF16 version of olmOCR-2-7B-1025-FP8. We recommend using the FP8 version for all practical purposes except further fine tuning.

Multimodal 7.0B ↓ 122.2K
Mage-VL logo
Mage-VL
microsoft

Mage-VL An Efficient Codec-Native Streaming Multimodal Foundation Model

Multimodal 4.0B ↓ 117.8K
InternVL3-1B-hf logo
InternVL3-1B-hf
OpenGVLab

InternVL3-1B Transformers 🤗 Implementation

Multimodal 0.94B ↓ 112.2K
InternVL3-8B logo
InternVL3-8B
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 7.94B ↓ 110.4K
OLMoE-1B-7B-0924 logo
OLMoE-1B-7B-0924
allenai

OLMoE-1B-7B is a Mixture-of-Experts LLM with 1B active and 7B total parameters released in September 2024 (0924). It yields state-of-the-art performance among models with a similar cost (1B) and is competitive with much larger models like Llama2-13B. OLMoE is 100% open-source.

Open Source 7.0B ↓ 110.4K
GLM-4.6V-Flash logo
GLM-4.6V-Flash
zai-org

This model is part of the GLM-V family of models, introduced in the paper GLM-4.1V-Thinking and GLM-4.5V: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

Open Source ↓ 108.3K
paligemma-3b-ft-cococap-448 logo
paligemma-3b-ft-cococap-448
google

Multimodal 3.0B ↓ 106.3K
NVIDIA-Nemotron-Parse-2.0 logo
NVIDIA-Nemotron-Parse-2.0
nvidia

Description: NVIDIA Nemotron Parse 2.0 transforms document images into structured, machine-readable representations with text, layout classes, bounding boxes, and reading-order information. Given a Red, Green, Blue (RGB) document image and a task prompt, the model produces format…

Multimodal 0.905B ↓ 101.7K
InternVL3-1B logo
InternVL3-1B
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 0.94B ↓ 99.9K
olmOCR-2-7B-1025-FP8 logo
olmOCR-2-7B-1025-FP8
allenai

Quantized to FP8 Version of olmOCR-2-7B-1025, using llmcompressor.

Multimodal 7.0B ↓ 99K
Cosmos-Reason1-7B logo
Cosmos-Reason1-7B
nvidia

Cosmos-Reason1: Physical AI Common Sense and Embodied Reasoning Models

Multimodal 7.0B ↓ 93.9K