AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,302 models for "Transformer" Compare
🤖
Qwen3-1.7B-Base
Qwen

Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Building upon extensive advancements in training data, model architecture, and optimization techniques, Qwen3 delivers the followin…

Open Source 1.7B ↓ 1.9M
🤖
Qwen2.5-Coder-14B-Instruct-AWQ
Qwen

Qwen2.5-Coder is the latest series of Code-Specific Qwen large language models (formerly known as CodeQwen). As of now, Qwen2.5-Coder has covered six mainstream model sizes, 0.5, 1.5, 3, 7, 14, 32 billion parameters, to meet the needs of different developers. Qwen2.5-Coder brings…

Code 14.7B ↓ 1.9M
🤖
DeepSeek-R1
deepseek-ai

We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With R…

Reasoning ↓ 1.9M
🤖
Qwen3-4B-Base
Qwen

Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Building upon extensive advancements in training data, model architecture, and optimization techniques, Qwen3 delivers the followin…

Open Source 4.0B ↓ 1.9M
🤖
Qwen2-VL-2B-Instruct
Qwen

We're excited to unveil Qwen2-VL , the latest iteration of our Qwen-VL model, representing nearly a year of innovation.

Multimodal 2.0B ↓ 1.8M
🤖
Qwen3-14B
Qwen

Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities,…

Open Source 14.8B ↓ 1.8M
🤖
Qwen3-8B-AWQ
Qwen

Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities,…

Open Source 8.0B ↓ 1.8M
🤖
Qwen2.5-0.5B
Qwen

Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2:

Open Source 0.5B ↓ 1.7M
🤖
Qwen2.5-Coder-32B-Instruct
Qwen

Qwen2.5-Coder is the latest series of Code-Specific Qwen large language models (formerly known as CodeQwen). As of now, Qwen2.5-Coder has covered six mainstream model sizes, 0.5, 1.5, 3, 7, 14, 32 billion parameters, to meet the needs of different developers. Qwen2.5-Coder brings…

Code 32.5B ↓ 1.7M
🤖
Qwen2-VL-7B-Instruct-AWQ
Qwen

We're excited to unveil Qwen2-VL , the latest iteration of our Qwen-VL model, representing nearly a year of innovation.

Multimodal 7.0B ↓ 1.7M
🤖
Qwen2.5-VL-7B-Instruct-AWQ
Qwen

In the past five months since Qwen2-VL’s release, numerous developers have built new models on the Qwen2-VL vision-language models, providing us with valuable feedback. During this period, we focused on building more useful vision-language models. Today, we are excited to introdu…

Multimodal 7.0B ↓ 1.6M
🤖
Meta-Llama-3-8B-Instruct
meta-llama

Open Source 8.0B ↓ 1.6M