LLM Models
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
[\[📂 GitHub\]](https://github.com/bytedance/SAIL) [\[📜 paper\]](https://arxiv.org/abs/2504.10462) [\[🚀 Quick Start\]]( quick-start)
Our Swallow-MX-8x7b-NVE-v0.1 model has undergone continuous pre-training from the Mixtral-8x7B-Instruct-v0.1, primarily with the addition of Japanese language data.
Multimodal Markup Document Models (MarkupDM)
Llama 3 Youko 70B Instruct (rinna/llama-3-youko-70b-instruct)
Llama3 Swallow - Built with Meta Llama 3
This is a Japanese continually pre-trained model based on meta-llama/Meta-Llama-3.1-70B-Instruct.
UI-TARS-72B-SFT UI-TARS-2B-SFT UI-TARS-7B-SFT UI-TARS-7B-DPO (Recommended) UI-TARS-72B-SFT UI-TARS-72B-DPO (Recommended) Introduction
Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.
Qwen3-Swallow v0.2 is a family of large language models available in 8B , 30B-A3B , and 32B parameter sizes. Built as bilingual Japanese-English models, they were developed through Continual Pre-Training (CPT), Supervised Fine-Tuning (SFT), and Reinforcement Learning with Verifia…