AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,294 models for "Transformer" Compare
🤖
SAIL-7B
ByteDance-Seed

[\[📂 GitHub\]](https://github.com/bytedance/SAIL) [\[📜 paper\]](https://arxiv.org/abs/2504.10462) [\[🚀 Quick Start\]]( quick-start)

Open Source 7.0B ↓ 73
🤖
Swallow-MX-8x7b-NVE-v0.1
tokyotech-llm

Our Swallow-MX-8x7b-NVE-v0.1 model has undergone continuous pre-training from the Mixtral-8x7B-Instruct-v0.1, primarily with the addition of Japanese language data.

Open Source 7.0B ↓ 72
🤖
llama-3-youko-70b-instruct
rinna

Llama 3 Youko 70B Instruct (rinna/llama-3-youko-70b-instruct)

Open Source 70.0B ↓ 72
🤖
Llama-3-Swallow-70B-Instruct-v0.1
tokyotech-llm

Llama3 Swallow - Built with Meta Llama 3

Open Source 70.0B ↓ 71
🤖
Llama-3.1-70B-Japanese-Instruct-2407
cyberagent

This is a Japanese continually pre-trained model based on meta-llama/Meta-Llama-3.1-70B-Instruct.

Open Source 70.0B ↓ 70
🤖
Swallow-7b-NVE-hf
tokyotech-llm

Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.

Open Source 7.0B ↓ 68
🤖
Llama-3.1-8B-Instruct-MoAA-SFT
togethercomputer

This is the SFT model in our Mixture of Agents Alignment (MoAA) pipeline. This model is tuned on the Llama-3.1-8b-Instruct. MoAA is an approach that leverages collective intelligence from open‑source LLMs to advance alignment.

Open Source 8.0B ↓ 65
🤖
sarashina1-7b
sbintuitions

This repository provides Japanese language models trained by SB Intuitions.

Open Source 7.0B ↓ 65
🤖
bloomz-7b1-p3
bigscience

1. Model Summary 2. Use 3. Limitations 4. Training 5. Evaluation 7. Citation

Open Source ↓ 65
🤖
Swallow-70b-NVE-instruct-hf
tokyotech-llm

Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.

Open Source 70.0B ↓ 64
🤖
ERNIE-4.5-VL-424B-A47B-Base-Paddle
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Multimodal 424.0B ↓ 64
🤖
Swallow-7b-plus-hf
tokyotech-llm

Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.

Open Source 7.0B ↓ 62