AI Agent Hub

LLM Models · tokyotech-llm

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

61 models from tokyotech-llm Compare
Qwen3-Swallow-8B-SFT-v0.2 logo
Qwen3-Swallow-8B-SFT-v0.2
tokyotech-llm

Qwen3-Swallow v0.2 is a family of large language models available in 8B , 30B-A3B , and 32B parameter sizes. Built as bilingual Japanese-English models, they were developed through Continual Pre-Training (CPT), Supervised Fine-Tuning (SFT), and Reinforcement Learning with Verifia…

open-source 8.0B ↓ 1.7K
GPT-OSS-Swallow-120B-RL-v0.1 logo
GPT-OSS-Swallow-120B-RL-v0.1
tokyotech-llm

GPT-OSS-Swallow v0.1 is a family of large language models available in 20B and 120B parameter sizes. Built as bilingual Japanese-English models, they were developed through Continual Pre-Training (CPT), Supervised Fine-Tuning (SFT), and Reinforcement Learning with Verifiable Rewa…

open-source 120.0B ↓ 1.2K
Gemma-2-Llama-Swallow-9b-pt-v0.1 logo
Gemma-2-Llama-Swallow-9b-pt-v0.1
tokyotech-llm

Gemma-2-Llama-Swallow series was built by continual pre-training on the gemma-2 models. Gemma 2 Swallow enhanced the Japanese language capabilities of the original Gemma 2 while retaining the English language capabilities. We use approximately 200 billion tokens that were sampled…

open-source 9.0B ↓ 1.1K
Qwen3-Swallow-8B-RL-v0.2-AWQ-INT4 logo
Qwen3-Swallow-8B-RL-v0.2-AWQ-INT4
tokyotech-llm

Qwen3-Swallow v0.2 is a family of large language models available in 8B , 30B-A3B , and 32B parameter sizes. Built as bilingual Japanese-English models, they were developed through Continual Pre-Training (CPT), Supervised Fine-Tuning (SFT), and Reinforcement Learning with Verifia…

open-source 8.0B ↓ 1K
GPT-OSS-Swallow-20B-SFT-v0.1 logo
GPT-OSS-Swallow-20B-SFT-v0.1
tokyotech-llm

GPT-OSS-Swallow v0.1 is a family of large language models available in 20B and 120B parameter sizes. Built as bilingual Japanese-English models, they were developed through Continual Pre-Training (CPT), Supervised Fine-Tuning (SFT), and Reinforcement Learning with Verifiable Rewa…

open-source 20.0B ↓ 977
Gemma-2-Llama-Swallow-9b-it-v0.1 logo
Gemma-2-Llama-Swallow-9b-it-v0.1
tokyotech-llm

Gemma-2-Llama-Swallow series was built by continual pre-training on the gemma-2 models. Gemma 2 Swallow enhanced the Japanese language capabilities of the original Gemma 2 while retaining the English language capabilities. We use approximately 200 billion tokens that were sampled…

open-source 9.0B ↓ 751
GPT-OSS-Swallow-120B-RL-v0.1-MXFP4 logo
GPT-OSS-Swallow-120B-RL-v0.1-MXFP4
tokyotech-llm

GPT-OSS-Swallow v0.1 is a family of large language models available in 20B and 120B parameter sizes. Built as bilingual Japanese-English models, they were developed through Continual Pre-Training (CPT), Supervised Fine-Tuning (SFT), and Reinforcement Learning with Verifiable Rewa…

open-source 120.0B ↓ 744
Gemma-2-Llama-Swallow-2b-it-v0.1 logo
Gemma-2-Llama-Swallow-2b-it-v0.1
tokyotech-llm

Gemma-2-Llama-Swallow series was built by continual pre-training on the gemma-2 models. Gemma 2 Swallow enhanced the Japanese language capabilities of the original Gemma 2 while retaining the English language capabilities. We use approximately 200 billion tokens that were sampled…

open-source 2.0B ↓ 743
Qwen3-Swallow-30B-A3B-SFT-v0.2 logo
Qwen3-Swallow-30B-A3B-SFT-v0.2
tokyotech-llm

Qwen3-Swallow v0.2 is a family of large language models available in 8B , 30B-A3B , and 32B parameter sizes. Built as bilingual Japanese-English models, they were developed through Continual Pre-Training (CPT), Supervised Fine-Tuning (SFT), and Reinforcement Learning with Verifia…

open-source 30.0B ↓ 682
Llama-3.1-Swallow-8B-Instruct-v0.1 logo
Llama-3.1-Swallow-8B-Instruct-v0.1
tokyotech-llm

Llama 3.1 Swallow is a series of large language models (8B, 70B) that were built by continual pre-training on the Meta Llama 3.1 models. Llama 3.1 Swallow enhanced the Japanese language capabilities of the original Llama 3.1 while retaining the English language capabilities. We u…

open-source 8.0B ↓ 652
Qwen3-Swallow-30B-A3B-RL-v0.2 logo
Qwen3-Swallow-30B-A3B-RL-v0.2
tokyotech-llm

Qwen3-Swallow v0.2 is a family of large language models available in 8B , 30B-A3B , and 32B parameter sizes. Built as bilingual Japanese-English models, they were developed through Continual Pre-Training (CPT), Supervised Fine-Tuning (SFT), and Reinforcement Learning with Verifia…

open-source 30.0B ↓ 625
Qwen3-Swallow-30B-A3B-RL-v0.2-AWQ-INT4 logo
Qwen3-Swallow-30B-A3B-RL-v0.2-AWQ-INT4
tokyotech-llm

Qwen3-Swallow v0.2 is a family of large language models available in 8B , 30B-A3B , and 32B parameter sizes. Built as bilingual Japanese-English models, they were developed through Continual Pre-Training (CPT), Supervised Fine-Tuning (SFT), and Reinforcement Learning with Verifia…

open-source 30.0B ↓ 562