LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
[\[📂 GitHub\]](https://github.com/bytedance/SAIL) [\[📜 paper\]](https://arxiv.org/abs/2504.10462) [\[🚀 Quick Start\]]( quick-start)
Our Swallow-MX-8x7b-NVE-v0.1 model has undergone continuous pre-training from the Mixtral-8x7B-Instruct-v0.1, primarily with the addition of Japanese language data.
Llama 3 Youko 70B Instruct (rinna/llama-3-youko-70b-instruct)
Llama3 Swallow - Built with Meta Llama 3
This is a Japanese continually pre-trained model based on meta-llama/Meta-Llama-3.1-70B-Instruct.
Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.
This is the SFT model in our Mixture of Agents Alignment (MoAA) pipeline. This model is tuned on the Llama-3.1-8b-Instruct. MoAA is an approach that leverages collective intelligence from open‑source LLMs to advance alignment.
This repository provides Japanese language models trained by SB Intuitions.
1. Model Summary 2. Use 3. Limitations 4. Training 5. Evaluation 7. Citation
Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.
[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.
Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.