AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,877 models Compare
🤖
AHN-DN-for-Qwen-2.5-Instruct-14B
ByteDance-Seed

AHN: Artificial Hippocampus Networks for Efficient Long-Context Modeling

Open Source 14.0B ↓ 58
🤖
llama-3-youko-70b
rinna

Llama 3 Youko 70B (rinna/llama-3-youko-70b)

Open Source 70.0B ↓ 57
🤖
Swallow-13b-NVE-hf
tokyotech-llm

Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.

Open Source 13.0B ↓ 57
🤖
LongCat-Flash-Thinking-2601-FP8
meituan-longcat

We introduce an updated version of LongCat-Flash-Thinking-2601, a powerful and efficient Large Reasoning Model (LRM) with 560 billion total parameters, built upon an innovative Mixture-of-Experts (MoE) architecture. Beyond inheriting the domain-parallel training recipe in our pre…

Open Source ↓ 55
🤖
AHN-GDN-for-Qwen-2.5-Instruct-7B
ByteDance-Seed

AHN: Artificial Hippocampus Networks for Efficient Long-Context Modeling

Open Source 7.0B ↓ 55
🤖
AHN-DN-for-Qwen-2.5-Instruct-7B
ByteDance-Seed

AHN: Artificial Hippocampus Networks for Efficient Long-Context Modeling

Open Source 7.0B ↓ 54
🤖
ERNIE-4.5-VL-28B-A3B-Base-PT
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Multimodal 28.0B ↓ 54
🤖
Gemma-2-Llama-Swallow-27b-pt-v0.1
tokyotech-llm

Gemma-2-Llama-Swallow series was built by continual pre-training on the gemma-2 models. Gemma 2 Swallow enhanced the Japanese language capabilities of the original Gemma 2 while retaining the English language capabilities. We use approximately 200 billion tokens that were sampled…

Open Source 27.0B ↓ 51
🤖
c4ai-command-r-plus-4bit
CohereLabs

Open Source ↓ 51
🤖
sarashina1-13b
sbintuitions

This repository provides Japanese language models trained by SB Intuitions.

Open Source 13.0B ↓ 50
🤖
ERNIE-4.5-21B-A3B-Paddle
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Open Source 21.0B ↓ 50
🤖
sarashina1-65b
sbintuitions

This repository provides Japanese language models trained by SB Intuitions.

Open Source 65.0B ↓ 49