AI Agent Hub

LLM Models · tokyotech-llm

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

61 models from tokyotech-llm Compare
Swallow-70b-instruct-v0.1 logo
Swallow-70b-instruct-v0.1
tokyotech-llm

Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.

open-source 70.0B ↓ 74
Swallow-MX-8x7b-NVE-v0.1 logo
Swallow-MX-8x7b-NVE-v0.1
tokyotech-llm

Our Swallow-MX-8x7b-NVE-v0.1 model has undergone continuous pre-training from the Mixtral-8x7B-Instruct-v0.1, primarily with the addition of Japanese language data.

open-source 7.0B ↓ 72
Llama-3-Swallow-70B-Instruct-v0.1 logo
Llama-3-Swallow-70B-Instruct-v0.1
tokyotech-llm

Llama3 Swallow - Built with Meta Llama 3

open-source 70.0B ↓ 71
Swallow-7b-NVE-hf logo
Swallow-7b-NVE-hf
tokyotech-llm

Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.

open-source 7.0B ↓ 68
Qwen3-Swallow-30B-A3B-CPT-v0.2 logo
Qwen3-Swallow-30B-A3B-CPT-v0.2
tokyotech-llm

Qwen3-Swallow v0.2 is a family of large language models available in 8B , 30B-A3B , and 32B parameter sizes. Built as bilingual Japanese-English models, they were developed through Continual Pre-Training (CPT), Supervised Fine-Tuning (SFT), and Reinforcement Learning with Verifia…

open-source 30.0B ↓ 67
Swallow-70b-NVE-instruct-hf logo
Swallow-70b-NVE-instruct-hf
tokyotech-llm

Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.

open-source 70.0B ↓ 64
Llama-3.1-Swallow-70B-v0.1 logo
Llama-3.1-Swallow-70B-v0.1
tokyotech-llm

Llama 3.1 Swallow is a series of large language models (8B, 70B) that were built by continual pre-training on the Meta Llama 3.1 models. Llama 3.1 Swallow enhanced the Japanese language capabilities of the original Llama 3.1 while retaining the English language capabilities. We u…

open-source 70.0B ↓ 63
Swallow-7b-plus-hf logo
Swallow-7b-plus-hf
tokyotech-llm

Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.

open-source 7.0B ↓ 62
Llama-3-Swallow-70B-v0.1 logo
Llama-3-Swallow-70B-v0.1
tokyotech-llm

Llama3 Swallow - Built with Meta Llama 3

open-source 70.0B ↓ 61
Swallow-7b-NVE-instruct-hf logo
Swallow-7b-NVE-instruct-hf
tokyotech-llm

Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.

open-source 7.0B ↓ 60
Swallow-70b-hf logo
Swallow-70b-hf
tokyotech-llm

Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.

open-source 70.0B ↓ 58
Swallow-13b-NVE-hf logo
Swallow-13b-NVE-hf
tokyotech-llm

Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.

open-source 13.0B ↓ 57