AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,315 models for "Transformer" Compare
🤖
Llama-xLAM-2-8b-fc-r
Salesforce

Large Action Models (LAMs) are advanced language models designed to enhance decision-making by translating user intentions into executable actions. As the brains of AI agents , LAMs autonomously plan and execute tasks to achieve specific goals, making them invaluable for automati…

Open Source 8.0B ↓ 20.2K
🤖
granite-4.0-h-micro
ibm-granite

📣 Update [10-07-2025]: Added a default system prompt to the chat template to guide the model towards more professional, accurate, and safe responses.

Open Source 3.0B ↓ 20.1K
🤖
internlm2-chat-20b
internlm

💻Github Repo • 🤔Reporting Issues • 📜Technical Report

Open Source 20.0B ↓ 19.8K
🤖
internlm2-20b
internlm

💻Github Repo • 🤔Reporting Issues • 📜Technical Report

Open Source 20.0B ↓ 19.7K
🤖
internlm2-base-7b
internlm

💻Github Repo • 🤔Reporting Issues • 📜Technical Report

Open Source 7.0B ↓ 19.7K
🤖
LFM2.5-1.2B-JP-202606
LiquidAI

LFM2.5-1.2B-JP-202606 is our latest general purpose Japanese chat model, delivering significant improvements in knowledge, instruction following, math, code, and tool-use over both the models of comparable size and LFM2.5-1.2B-JP. It sets a new benchmark for state-of-the-art perf…

Open Source 1.2B ↓ 19.6K
🤖
Nous-Hermes-2-Mixtral-8x7B-DPO
NousResearch

Nous Hermes 2 Mixtral 8x7B DPO is the new flagship Nous Research model trained over the Mixtral 8x7B MoE LLM.

Open Source 7.0B ↓ 19.6K
🤖
ERNIE-4.5-0.3B-PT
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Open Source 0.36B ↓ 19.5K
🤖
Intern-S1
internlm

💻Github Repo • 🤗Model Collections • 📜Technical Report • 💬Online Chat

Open Source ↓ 19.5K
🤖
Moonlight-16B-A3B
moonshotai

Tech Report HuggingFace Megatron(coming soon)

Open Source 16.0B ↓ 19.3K
🤖
MiniMax-VL-01
MiniMaxAI

1. Introduction We are delighted to introduce our MiniMax-VL-01 model. It adopts the "ViT-MLP-LLM" framework, which is a commonly used technique in the field of multimodal large language models. The model is initialized and trained with three key parts: a 303-million-parameter Vi…

Multimodal 456.0B ↓ 18.9K
🤖
Llama-2-13b-hf
meta-llama

Open Source 13.0B ↓ 18.4K