AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

440 models for "Base Model" Compare
Meta-Llama-3-8B-Alternate-Tokenizer logo
Meta-Llama-3-8B-Alternate-Tokenizer
NousResearch

An alternate Meta-Llama-3-8B Repo for the Hermes Tokenizer

Open Source 8.0B ↓ 177
OREAL-7B-SFT logo
OREAL-7B-SFT
internlm

--- license: apache-2.0 library name: transformers base model: - Qwen/Qwen2.5-7B pipeline tag: text-generation ---

Open Source 7.0B ↓ 176
AquilaDense-16B logo
AquilaDense-16B
BAAI

AquilaMoE: Efficient Training for MoE Models with Scale-Up and Scale-Out Strategies Language Foundation Model & Software Team Beijing Academy of Artificial Intelligence (BAAI) [Paper(released soon)] [Code] [github]

Open Source 16.0B ↓ 176
bloom-7b1-intermediate logo
bloom-7b1-intermediate
bigscience

WARNING: The checkpoints on this repo are not fully trained model. Evaluations of intermediary checkpoints and the final model will be added when conducted (see below).

Open Source ↓ 172
AHN-GDN-for-Qwen-2.5-Instruct-3B logo
AHN-GDN-for-Qwen-2.5-Instruct-3B
ByteDance-Seed

AHN: Artificial Hippocampus Networks for Efficient Long-Context Modeling

Open Source 3.0B ↓ 171
Swallow-70b-NVE-hf logo
Swallow-70b-NVE-hf
tokyotech-llm

Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.

Open Source 70.0B ↓ 170
smollm-360M-instruct-add-basics logo
smollm-360M-instruct-add-basics
HuggingFaceTB

Model Summary Chat with the model at: https://huggingface.co/spaces/HuggingFaceTB/instant-smol

Open Source ↓ 169
AHN-Mamba2-for-Qwen-2.5-Instruct-14B logo
AHN-Mamba2-for-Qwen-2.5-Instruct-14B
ByteDance-Seed

AHN: Artificial Hippocampus Networks for Efficient Long-Context Modeling

Open Source 14.0B ↓ 168
ERNIE-4.5-0.3B-Base-Paddle logo
ERNIE-4.5-0.3B-Base-Paddle
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Open Source 0.3B ↓ 167
ERNIE-4.5-VL-28B-A3B-Base-Paddle logo
ERNIE-4.5-VL-28B-A3B-Base-Paddle
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Multimodal 28.0B ↓ 160
AHN-Mamba2-for-Qwen-2.5-Instruct-3B logo
AHN-Mamba2-for-Qwen-2.5-Instruct-3B
ByteDance-Seed

AHN: Artificial Hippocampus Networks for Efficient Long-Context Modeling

Open Source 3.0B ↓ 158
Baichuan2-13B-Chat-4bits logo
Baichuan2-13B-Chat-4bits
baichuan-inc

🦉GitHub 💬WeChat 百川API支持搜索增强和192K长窗口,新增百川搜索增强知识库、限时免费! 🚀 百川大模型在线对话平台 已正式向公众开放 🎉

Open Source 13.0B ↓ 152