AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,297 models for "Transformer" Compare
🤖
Llama-SEA-LION-v3-70B
aisingapore

Llama-SEA-LION-v3-70B SEA-LION is a collection of Large Language Models (LLMs) which have been pretrained and instruct-tuned for the Southeast Asia (SEA) region.

Open Source 70.0B ↓ 242
🤖
mt0-xxl
bigscience

1. Model Summary 2. Use 3. Limitations 4. Training 5. Evaluation 7. Citation

Open Source ↓ 242
🤖
AquilaMoE
BAAI

AquilaMoE: Efficient Training for MoE Models with Scale-Up and Scale-Out Strategies Language Foundation Model & Software Team Beijing Academy of Artificial Intelligence (BAAI) [Paper(released soon)] [Code] [github]

Open Source ↓ 241
🤖
ERNIE-4.5-VL-424B-A47B-Base-PT
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Multimodal 424.0B ↓ 239
🤖
JanusCoderV-7B
internlm

💻Github Repo • 🤗Model Collections • 📜Technical Report

Code 7.0B ↓ 235
🤖
AquilaMed-RL
BAAI

Aquila is a large language model independently developed by BAAI. Building upon the Aquila model, we continued pre-training, SFT (Supervised Fine-Tuning), and RL (Reinforcement Learning) through a multi-stage training process, ultimately resulting in the AquilaMed-RL model. This…

Open Source ↓ 233
🤖
Spatial-SSRL-7B
internlm

📖 Paper 🏠 Github 🤗 Spatial-SSRL-7B Model 🤗 Spatial-SSRL-3B Model 🤗 Spatial-SSRL-Qwen3VL-4B Model 🤗 Spatial-SSRL-81k Dataset 📰 Daily Paper

Open Source 7.0B ↓ 232
🤖
llama-3-youko-8b
rinna

Llama 3 Youko 8B (rinna/llama-3-youko-8b)

Open Source 8.0B ↓ 230
🤖
ERNIE-4.5-0.3B-Paddle
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Open Source 0.3B ↓ 228
🤖
japanese-stablelm-3b-4e1t-instruct
stabilityai

This is a 3B-parameter decoder-only Japanese language model fine-tuned on instruction-following datasets, built on top of the base model Japanese StableLM-3B-4E1T Base.

Open Source 3.0B ↓ 227
🤖
ERNIE-4.5-21B-A3B-Base-PT
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Open Source 21.0B ↓ 226
🤖
japanese-stablelm-base-beta-70b
stabilityai

A cute robot wearing a kimono writes calligraphy with one single brush — Stable Diffusion XL

Open Source 70.0B ↓ 225