AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,302 models for "Transformer" Compare
🤖
gemma-3-4b-it
google

Multimodal 4.0B ↓ 1.5M
🤖
SmolVLM2-500M-Video-Instruct
HuggingFaceTB

SmolVLM2-500M-Video is a lightweight multimodal model designed to analyze video content. The model processes videos, images, and text inputs to generate text outputs - whether answering questions about media files, comparing visual content, or transcribing text from images. Despi…

Multimodal 0.5B ↓ 1.5M
🤖
DeepSeek-V3.2
deepseek-ai

DeepSeek-V3.2: Efficient Reasoning & Agentic AI

Open Source ↓ 1.5M
🤖
Qwen2.5-VL-32B-Instruct
Qwen

Latest Updates: In addition to the original formula, we have further enhanced Qwen2.5-VL-32B's mathematical and problem-solving abilities through reinforcement learning. This has also significantly improved the model's subjective user experience, with response styles adjusted to…

Multimodal 32.0B ↓ 1.5M
🤖
phi-2
microsoft

Phi-2 is a Transformer with 2.7 billion parameters. It was trained using the same data sources as Phi-1.5, augmented with a new data source that consists of various NLP synthetic texts and filtered websites (for safety and educational value). When assessed against benchmarks test…

Open Source ↓ 1.5M
🤖
Llama-3.2-3B-Instruct
meta-llama

Open Source 3.21B ↓ 1.4M
🤖
SmolLM2-135M-Instruct
HuggingFaceTB

1. Model Summary 2. Limitations 3. Training 4. License 5. Citation

Open Source ↓ 1.4M
🤖
OpenELM-1_1B-Instruct
apple

Sachin Mehta, Mohammad Hossein Sekhavat, Qingqing Cao, Maxwell Horton, Yanzi Jin, Chenfan Sun, Iman Mirzadeh, Mahyar Najibi, Dmitry Belenko, Peter Zatloukal, Mohammad Rastegari

Open Source 1.1B ↓ 1.4M
🤖
Qwen2.5-Coder-32B-Instruct-AWQ
Qwen

Qwen2.5-Coder is the latest series of Code-Specific Qwen large language models (formerly known as CodeQwen). As of now, Qwen2.5-Coder has covered six mainstream model sizes, 0.5, 1.5, 3, 7, 14, 32 billion parameters, to meet the needs of different developers. Qwen2.5-Coder brings…

Code 32.5B ↓ 1.4M
🤖
Qwen3-30B-A3B-Instruct-2507
Qwen

We introduce the updated version of the Qwen3-30B-A3B non-thinking mode , named Qwen3-30B-A3B-Instruct-2507 , featuring the following key enhancements:

Open Source 30.0B ↓ 1.4M
🤖
Llama-3.2-1B
meta-llama

Open Source 1.23B ↓ 1.3M
🤖
Qwen2-VL-7B-Instruct
Qwen

We're excited to unveil Qwen2-VL , the latest iteration of our Qwen-VL model, representing nearly a year of innovation.

Multimodal 7.0B ↓ 1.3M