AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,305 models for "Transformer" Compare
🤖
deepseek-coder-6.7b-instruct
deepseek-ai

[🏠Homepage] [🤖 Chat with DeepSeek Coder] [Discord] [Wechat(微信)]

Code 6.7B ↓ 396.1K
🤖
DeepSeek-R1-Distill-Llama-8B
deepseek-ai

We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With R…

Reasoning 8.0B ↓ 388.6K
🤖
Llama-3.2-3B
meta-llama

Open Source 3.21B ↓ 378.2K
🤖
DeepSeek-R1-Distill-Qwen-7B
deepseek-ai

We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With R…

Reasoning 7.0B ↓ 374K
🤖
EXAONE-3.5-7.8B-Instruct
LGAI-EXAONE

We introduce EXAONE 3.5, a collection of instruction-tuned bilingual (English and Korean) generative models ranging from 2.4B to 32B parameters, developed and released by LG AI Research. EXAONE 3.5 language models include: 1) 2.4B model optimized for deployment on small or resour…

Open Source 7.8B ↓ 369.4K
🤖
DeepSeek-V2-Lite
deepseek-ai

Model Download Evaluation Results Model Architecture API Platform License Citation

Open Source ↓ 368.6K
🤖
MiniCPM-V-4_5
openbmb

A GPT-4o Level MLLM for Single Image, Multi Image and High-FPS Video Understanding on Your Phone

Multimodal 8.7B ↓ 359.1K
🤖
SmolLM2-1.7B
HuggingFaceTB

1. Model Summary 2. Evaluation 3. Limitations 4. Training 5. License 6. Citation

Open Source 1.7B ↓ 347.4K
🤖
Llama-2-7b-chat-hf
meta-llama

Open Source 7.0B ↓ 343.8K
🤖
EXAONE-3.5-7.8B-Instruct-AWQ
LGAI-EXAONE

We introduce EXAONE 3.5, a collection of instruction-tuned bilingual (English and Korean) generative models ranging from 2.4B to 32B parameters, developed and released by LG AI Research. EXAONE 3.5 language models include: 1) 2.4B model optimized for deployment on small or resour…

Open Source 7.8B ↓ 342K
🤖
Mistral-7B-v0.1
mistralai

The Mistral-7B-v0.1 Large Language Model (LLM) is a pretrained generative text model with 7 billion parameters. Mistral-7B-v0.1 outperforms Llama 2 13B on all benchmarks we tested.

Open Source 7.0B ↓ 336.9K
🤖
Kimi-VL-A3B-Instruct
moonshotai

📄 Tech Report     📄 Github     💬 Chat Web

Multimodal 16.0B ↓ 315.8K