AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,308 models for "Transformer" Compare
🤖
OLMoE-1B-7B-0125-Instruct
allenai

OLMoE-1B-7B-0125-Instruct January 2025 is post-trained variant of the OLMoE-1B-7B January 2025 model, which has undergone supervised finetuning on an OLMo-specific variant of the Tülu 3 dataset and further DPO training on this dataset, and finally RLVR training using this data. T…

Open Source 1.0B ↓ 264.8K
🤖
medgemma-1.5-4b-it
google

Multimodal 4.0B ↓ 263.8K
🤖
olmOCR-2-7B-1025-FP8
allenai

Quantized to FP8 Version of olmOCR-2-7B-1025, using llmcompressor.

Multimodal 7.0B ↓ 260K
🤖
MiniMax-M2
MiniMaxAI

Join Our 💬 WeChat 🧩 Discord community. MiniMax Agent ⚡️ API (Now Free for a limited time!) MCP MiniMax Website 🤗 Hugging Face 🐙 GitHub 🤖️ ModelScope 📄 License: MIT

Open Source ↓ 258.3K
🤖
paligemma-3b-ft-cococap-448
google

Multimodal 3.0B ↓ 253.1K
🤖
InternVL2_5-4B
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 Mini-InternVL\]](https://arxiv.org/abs/2410.16261) [\[📜 InternVL 2.5\]](https://huggingface.co/…

Multimodal 3.7B ↓ 230.9K
🤖
Florence-2-base-ft
microsoft

Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks

Multimodal 0.23B ↓ 230.5K
🤖
tiny_starcoder_py
bigcode

This is a 164M parameters model with the same architecture as StarCoder (8k context length, MQA & FIM). It was trained on the Python data from StarCoderData for ~6 epochs which amounts to 100B tokens.

Code 0.164B ↓ 228.7K
🤖
DeepSeek-V3.1
deepseek-ai

DeepSeek-V3.1 is a hybrid model that supports both thinking mode and non-thinking mode. Compared to the previous version, this upgrade brings improvements in multiple aspects:

Open Source ↓ 228.6K
🤖
LLaDA2.0-mini
inclusionAI

LLaDA2.0-mini is a diffusion language model featuring a 16BA1B Mixture-of-Experts (MoE) architecture. As an enhanced, instruction-tuned iteration of the LLaDA series, it is optimized for practical applications.

Open Source ↓ 223.3K
🤖
GLM-4.1V-9B-Thinking
zai-org

📖 View the GLM-4.1V-9B-Thinking paper . 📍 Using GLM-4.1V-9B-Thinking API at Zhipu Foundation Model Open Platform

Open Source 9.0B ↓ 220.4K
🤖
Llama-3.1-Nemotron-Nano-8B-v1
nvidia

Llama-3.1-Nemotron-Nano-8B-v1 is a large language model (LLM) which is a derivative of Meta Llama-3.1-8B-Instruct (AKA the reference model). It is a reasoning model that is post trained for reasoning, human chat preferences, and tasks, such as RAG and tool calling.

Reasoning 8.0B ↓ 218.1K