LLM Models · tokyotech-llm
Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.
Medical-Qwen3-Swallow-8B is a medical-domain language model based on tokyotech-llm/Qwen3-Swallow-8B-RL-v0.2. It is designed to support research and development toward safe and trustworthy AI for Japanese clinical settings.
Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.
Gemma-2-Llama-Swallow series was built by continual pre-training on the gemma-2 models. Gemma 2 Swallow enhanced the Japanese language capabilities of the original Gemma 2 while retaining the English language capabilities. We use approximately 200 billion tokens that were sampled…
Llama 3.3 Swallow is a large language model (70B) that was built by continual pre-training on the Meta Llama 3.3 model. Llama 3.3 Swallow enhanced the Japanese language capabilities of the original Llama 3.3 while retaining the English language capabilities. We use approximately…
GPT-OSS-Swallow v0.1 is a family of large language models available in 20B and 120B parameter sizes. Built as bilingual Japanese-English models, they were developed through Continual Pre-Training (CPT), Supervised Fine-Tuning (SFT), and Reinforcement Learning with Verifiable Rewa…
Llama 3.1 Swallow is a series of large language models (8B, 70B) that were built by continual pre-training on the Meta Llama 3.1 models. Llama 3.1 Swallow enhanced the Japanese language capabilities of the original Llama 3.1 while retaining the English language capabilities. We u…
Medical-Qwen3-Swallow-30B-A3B is a medical-domain language model based on tokyotech-llm/Qwen3-Swallow-30B-A3B-RL-v0.2 . It is designed to support research and development toward safe and trustworthy AI for Japanese clinical settings.
Llama 3.1 Swallow is a series of large language models (8B, 70B) that were built by continual pre-training on the Meta Llama 3.1 models. Llama 3.1 Swallow enhanced the Japanese language capabilities of the original Llama 3.1 while retaining the English language capabilities. We u…
Llama 3.3 Swallow is a large language model (70B) that was built by continual pre-training on the Meta Llama 3.3 model. Llama 3.3 Swallow enhanced the Japanese language capabilities of the original Llama 3.3 while retaining the English language capabilities. We use approximately…
Medical-Qwen3-Swallow-32B is a medical-domain language model based on tokyotech-llm/Qwen3-Swallow-32B-RL-v0.2. It is designed to support research and development toward safe and trustworthy AI for Japanese clinical settings.
Medical-GPT-OSS-Swallow-120B is a medical-domain language model based on tokyotech-llm/GPT-OSS-Swallow-120B-RL-v0.1. It is designed to support research and development toward safe and trustworthy AI for Japanese clinical settings.
Qwen3-Swallow v0.2 is a family of large language models available in 8B , 30B-A3B , and 32B parameter sizes. Built as bilingual Japanese-English models, they were developed through Continual Pre-Training (CPT), Supervised Fine-Tuning (SFT), and Reinforcement Learning with Verifia…