AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

440 models for "Base Model" Compare
SuperApriel-15b-Base logo
SuperApriel-15b-Base
ServiceNow-AI

A 15B-parameter token-mixer supernet derived from Apriel-1.6 via stochastic distillation. Every decoder layer exposes four trained mixer options —Full Attention, Sliding Window Attention, Gated DeltaNet, and Kimi Delta Attention—enabling flexible architecture selection from a sin…

Open Source 15.0B ↓ 150
BFS-Prover-V2-32B logo
BFS-Prover-V2-32B
ByteDance-Seed

BFS-Prover-V2: Scaling up Multi-Turn Off-Policy RL and Multi-Agent Tree Search for LLM Step-Provers

Open Source 32.0B ↓ 150
AHN-Mamba2-for-Qwen-2.5-Instruct-7B logo
AHN-Mamba2-for-Qwen-2.5-Instruct-7B
ByteDance-Seed

AHN: Artificial Hippocampus Networks for Efficient Long-Context Modeling

Open Source 7.0B ↓ 148
ERNIE-4.5-VL-424B-A47B-Base-Paddle logo
ERNIE-4.5-VL-424B-A47B-Base-Paddle
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Multimodal 424.0B ↓ 147
Swallow-7b-NVE-hf logo
Swallow-7b-NVE-hf
tokyotech-llm

Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.

Open Source 7.0B ↓ 143
AHN-DN-for-Qwen-2.5-Instruct-14B logo
AHN-DN-for-Qwen-2.5-Instruct-14B
ByteDance-Seed

AHN: Artificial Hippocampus Networks for Efficient Long-Context Modeling

Open Source 14.0B ↓ 142
Swallow-7b-NVE-instruct-hf logo
Swallow-7b-NVE-instruct-hf
tokyotech-llm

Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.

Open Source 7.0B ↓ 141
ERNIE-4.5-21B-A3B-Base-Paddle logo
ERNIE-4.5-21B-A3B-Base-Paddle
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Open Source 21.0B ↓ 140
Baichuan2-7B-Chat-4bits logo
Baichuan2-7B-Chat-4bits
baichuan-inc

🦉GitHub 💬WeChat 百川API支持搜索增强和192K长窗口,新增百川搜索增强知识库、限时免费! 🚀 百川大模型在线对话平台 已正式向公众开放 🎉

Open Source 7.0B ↓ 140
Swallow-70b-hf logo
Swallow-70b-hf
tokyotech-llm

Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.

Open Source 70.0B ↓ 134
AHN-GDN-for-Qwen-2.5-Instruct-7B logo
AHN-GDN-for-Qwen-2.5-Instruct-7B
ByteDance-Seed

AHN: Artificial Hippocampus Networks for Efficient Long-Context Modeling

Open Source 7.0B ↓ 134
Swallow-70b-NVE-instruct-hf logo
Swallow-70b-NVE-instruct-hf
tokyotech-llm

Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.

Open Source 70.0B ↓ 134