AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,304 models for "Transformer" Compare
🤖
Cosmos-Reason2-2B
nvidia

Multimodal 2.0B ↓ 982.2K
🤖
bloomz-560m
bigscience

1. Model Summary 2. Use 3. Limitations 4. Training 5. Evaluation 7. Citation

Open Source 0.56B ↓ 974.5K
🤖
Qwen2-1.5B-Instruct
Qwen

Qwen2 is the new series of Qwen large language models. For Qwen2, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters, including a Mixture-of-Experts model. This repo contains the instruction-tuned 1.5B Qwen2…

Open Source 1.5B ↓ 949.6K
🤖
gemma-3-12b-it
google

Multimodal 12.0B ↓ 908.4K
🤖
Qwen2-0.5B
Qwen

Qwen2 is the new series of Qwen large language models. For Qwen2, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters, including a Mixture-of-Experts model. This repo contains the 0.5B Qwen2 base language mod…

Open Source 0.5B ↓ 880.5K
🤖
Qwen2.5-Math-1.5B
Qwen

[!Warning] 🚨 Qwen2.5-Math mainly supports solving English and Chinese math problems through CoT and TIR. We do not recommend using this series of models for other tasks.

Reasoning 1.5B ↓ 859K
🤖
Qwen3-VL-235B-A22B-Instruct
Qwen

Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.

Multimodal 235.0B ↓ 825.5K
🤖
GLM-5-FP8
zai-org

👋 Join our WeChat or Discord community. 📖 Check out the GLM-5 Technical report . 📍 Use GLM-5 API services on Z.ai API Platform. 👉 One click to GLM-5 .

Open Source ↓ 825.1K
🤖
deepseek-coder-7b-instruct-v1.5
deepseek-ai

[🏠Homepage] [🤖 Chat with DeepSeek Coder] [Discord] [Wechat(微信)]

Code 7.0B ↓ 812.8K
🤖
Llama-2-7b-hf
meta-llama

Open Source 7.0B ↓ 812.4K
🤖
MiniMax-M2.5-NVFP4
nvidia

Description: The NVIDIA MiniMax-M2.5-NVFP4 model is the quantized version of MiniMax's MiniMax-M2.5 model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA MiniMax-M2.5 NVFP4 model is q…

Open Source ↓ 776.5K
🤖
Qwen3-Coder-30B-A3B-Instruct
Qwen

Qwen3-Coder is available in multiple sizes. Today, we're excited to introduce Qwen3-Coder-30B-A3B-Instruct . This streamlined model maintains impressive performance and efficiency, featuring the following key enhancements:

Code 30.0B ↓ 774.8K