AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,877 models Compare
🤖
Llama-3.1-8B-Instruct-MoAA-SFT
togethercomputer

This is the SFT model in our Mixture of Agents Alignment (MoAA) pipeline. This model is tuned on the Llama-3.1-8b-Instruct. MoAA is an approach that leverages collective intelligence from open‑source LLMs to advance alignment.

Open Source 8.0B ↓ 65
🤖
sarashina1-7b
sbintuitions

This repository provides Japanese language models trained by SB Intuitions.

Open Source 7.0B ↓ 65
🤖
bloomz-7b1-p3
bigscience

1. Model Summary 2. Use 3. Limitations 4. Training 5. Evaluation 7. Citation

Open Source ↓ 65
🤖
Swallow-70b-NVE-instruct-hf
tokyotech-llm

Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.

Open Source 70.0B ↓ 64
🤖
bloom-petals
bigscience

This model is a version of bigscience/bloom post-processed to be run at home using the Petals swarm.

Open Source ↓ 64
🤖
ERNIE-4.5-VL-424B-A47B-Base-Paddle
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Multimodal 424.0B ↓ 64
🤖
Llama-3.1-Swallow-70B-v0.1
tokyotech-llm

Llama 3.1 Swallow is a series of large language models (8B, 70B) that were built by continual pre-training on the Meta Llama 3.1 models. Llama 3.1 Swallow enhanced the Japanese language capabilities of the original Llama 3.1 while retaining the English language capabilities. We u…

Open Source 70.0B ↓ 63
🤖
Swallow-7b-plus-hf
tokyotech-llm

Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.

Open Source 7.0B ↓ 62
🤖
Llama-3.1-8B-Instruct-MoAA-DPO
togethercomputer

This is the DPO model in our Mixture of Agents Alignment (MoAA) pipeline. This model is tuned on the Llama-3.1-8b-Instruct. MoAA is an approach that leverages collective intelligence from open‑source LLMs to advance alignment.

Open Source 8.0B ↓ 61
🤖
gemma-2-9b-it-MoAA-DPO
togethercomputer

This is the DPO model in our Mixture of Agents Alignment (MoAA) pipeline. This model is tuned on the Gemma-2-9b-it. MoAA is an approach that leverages collective intelligence from open‑source LLMs to advance alignment.

Open Source 9.0B ↓ 61
🤖
Llama-3-Swallow-70B-v0.1
tokyotech-llm

Llama3 Swallow - Built with Meta Llama 3

Open Source 70.0B ↓ 61
🤖
gemma-2-9b-it-MoAA-SFT
togethercomputer

This is the SFT model in our Mixture of Agents Alignment (MoAA) pipeline. This model is tuned on the Gemma-2-9b-it. MoAA is an approach that leverages collective intelligence from open‑source LLMs to advance alignment.

Open Source 9.0B ↓ 61