AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

562 models for "Vision" Compare
InternVL3-14B logo
InternVL3-14B
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 15.1B ↓ 6.3K
deepseek-vl-1.3b-chat logo
deepseek-vl-1.3b-chat
deepseek-ai

Introducing DeepSeek-VL, an open-source Vision-Language (VL) Model designed for real-world vision and language understanding applications. DeepSeek-VL possesses general multimodal understanding capabilities, capable of processing logical diagrams, web pages, formula recognition,…

Multimodal 1.3B ↓ 5.9K
deepseek-vl-7b-chat logo
deepseek-vl-7b-chat
deepseek-ai

Introducing DeepSeek-VL, an open-source Vision-Language (VL) Model designed for real-world vision and language understanding applications. DeepSeek-VL possesses general multimodal understanding capabilities, capable of processing logical diagrams, web pages, formula recognition,…

Multimodal 7.0B ↓ 5.8K
FastVLM-0.5B logo
FastVLM-0.5B
apple

FastVLM: Efficient Vision Encoding for Vision Language Models

Multimodal 0.5B ↓ 5.8K
OLMo-2-0325-32B logo
OLMo-2-0325-32B
allenai

We introduce OLMo 2 32B, the largest model in the OLMo 2 family. OLMo 2 was pre-trained on OLMo-mix-1124 and uses Dolmino-mix-1124 for mid-training.

Open Source 32.0B ↓ 5.7K
granite-3.2-2b-instruct logo
granite-3.2-2b-instruct
ibm-granite

Model Summary: Granite-3.2-2B-Instruct is an 2-billion-parameter, long-context AI model fine-tuned for thinking capabilities. Built on top of Granite-3.1-2B-Instruct, it has been trained using a mix of permissively licensed open-source datasets and internally generated synthetic…

Reasoning 2.0B ↓ 5.4K
d1-3B logo
d1-3B
LiquidAI

d1-3B is a 3B parameter decision model built on LFM2.5-VL-3B. You give it a state (text, JSON, images, or a mix) and a set of questions. It returns calibrated, typed answers in one forward pass with zero output tokens .

Multimodal 3.12B ↓ 5.4K
InternVL3_5-2B-HF logo
InternVL3_5-2B-HF
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 2.3B ↓ 5.2K
OLMo-1B-0724-hf logo
OLMo-1B-0724-hf
allenai

OLMo 1B July 2024 is the latest version of the original OLMo 1B model rocking a 4.4 point increase in HellaSwag, among other evaluations improvements, from an improved version of the Dolma dataset and staged training. This version is for direct use with HuggingFace Transformers f…

Code 1.0B ↓ 4.9K
InternVL3-14B-hf logo
InternVL3-14B-hf
OpenGVLab

InternVL3-14B Transformers 🤗 Implementation

Multimodal 15.1B ↓ 4.8K
RedPajama-INCITE-Base-3B-v1 logo
RedPajama-INCITE-Base-3B-v1
togethercomputer

RedPajama-INCITE-Base-3B-v1 was developed by Together and leaders from the open-source AI community including Ontocord.ai, ETH DS3Lab, AAI CERC, Université de Montréal, MILA - Québec AI Institute, Stanford Center for Research on Foundation Models (CRFM), Stanford Hazy Research re…

Open Source 3.0B ↓ 4.8K
deepseek-vl2-small logo
deepseek-vl2-small
deepseek-ai

Introducing DeepSeek-VL2, an advanced series of large Mixture-of-Experts (MoE) Vision-Language Models that significantly improves upon its predecessor, DeepSeek-VL. DeepSeek-VL2 demonstrates superior capabilities across various tasks, including but not limited to visual question…

Multimodal 16.0B ↓ 4.5K