AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,108 models for "Chat" Compare
🤖
Swallow-70b-instruct-v0.1
tokyotech-llm

Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.

Open Source 70.0B ↓ 74
🤖
diafill-sarashina2.2-3b-instruct
sbintuitions

Model Summary DiaFill is a Japanese dialogue script generation model designed to produce natural, spoken-style dialogue scripts rich in fillers and brief utternaces. Unlike typical assistant models that respond to users, this model is fine-tuned to generate a multi-turn dialogue…

Open Source 3.0B ↓ 74
🤖
Apriel-H1-15b-Thinker-SFT
ServiceNow-AI

A 15B-parameter hybrid reasoning model combining Transformer attention and Mamba State Space layers for high efficiency and scalability. Derived from Apriel-Nemotron-15B-Thinker through progressive distillation, Apriel-H1 replaces less critical attention layers with linear Mamba…

Open Source 15.0B ↓ 73
🤖
llama-3-youko-70b-instruct
rinna

Llama 3 Youko 70B Instruct (rinna/llama-3-youko-70b-instruct)

Open Source 70.0B ↓ 72
🤖
Llama-3-Swallow-70B-Instruct-v0.1
tokyotech-llm

Llama3 Swallow - Built with Meta Llama 3

Open Source 70.0B ↓ 71
🤖
Llama-3.1-70B-Japanese-Instruct-2407
cyberagent

This is a Japanese continually pre-trained model based on meta-llama/Meta-Llama-3.1-70B-Instruct.

Open Source 70.0B ↓ 70
🤖
Qwen3-Swallow-30B-A3B-CPT-v0.2
tokyotech-llm

Qwen3-Swallow v0.2 is a family of large language models available in 8B , 30B-A3B , and 32B parameter sizes. Built as bilingual Japanese-English models, they were developed through Continual Pre-Training (CPT), Supervised Fine-Tuning (SFT), and Reinforcement Learning with Verifia…

Open Source 30.0B ↓ 67
🤖
Llama-3.1-8B-Instruct-MoAA-SFT
togethercomputer

This is the SFT model in our Mixture of Agents Alignment (MoAA) pipeline. This model is tuned on the Llama-3.1-8b-Instruct. MoAA is an approach that leverages collective intelligence from open‑source LLMs to advance alignment.

Open Source 8.0B ↓ 65
🤖
ERNIE-4.5-VL-424B-A47B-Base-Paddle
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Multimodal 424.0B ↓ 64
🤖
Llama-3-Swallow-70B-v0.1
tokyotech-llm

Llama3 Swallow - Built with Meta Llama 3

Open Source 70.0B ↓ 61
🤖
gemma-2-9b-it-MoAA-SFT
togethercomputer

This is the SFT model in our Mixture of Agents Alignment (MoAA) pipeline. This model is tuned on the Gemma-2-9b-it. MoAA is an approach that leverages collective intelligence from open‑source LLMs to advance alignment.

Open Source 9.0B ↓ 61
🤖
Qianfan-VL-8B
baidu

Qianfan-VL: Domain-Enhanced Universal Vision-Language Models

Multimodal 8.0B ↓ 60