AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,294 models for "Transformer" Compare
🤖
Llama-3.1-8B-Instruct-MoAA-DPO
togethercomputer

This is the DPO model in our Mixture of Agents Alignment (MoAA) pipeline. This model is tuned on the Llama-3.1-8b-Instruct. MoAA is an approach that leverages collective intelligence from open‑source LLMs to advance alignment.

Open Source 8.0B ↓ 61
🤖
gemma-2-9b-it-MoAA-DPO
togethercomputer

This is the DPO model in our Mixture of Agents Alignment (MoAA) pipeline. This model is tuned on the Gemma-2-9b-it. MoAA is an approach that leverages collective intelligence from open‑source LLMs to advance alignment.

Open Source 9.0B ↓ 61
🤖
gemma-2-9b-it-MoAA-SFT
togethercomputer

This is the SFT model in our Mixture of Agents Alignment (MoAA) pipeline. This model is tuned on the Gemma-2-9b-it. MoAA is an approach that leverages collective intelligence from open‑source LLMs to advance alignment.

Open Source 9.0B ↓ 61
🤖
Llama-3.1-8B-Dragonfly-Med-v2
togethercomputer

Note: Users are permitted to use this model in accordance with the Llama 3.1 Community License Agreement. Additionally, due to the licensing restrictions of the dataset used to train this model, which prohibits commercial use, the Dragonfly-Med model is restricted to non-commerci…

Open Source 8.0B ↓ 61
🤖
AHN-GDN-for-Qwen-2.5-Instruct-14B
ByteDance-Seed

AHN: Artificial Hippocampus Networks for Efficient Long-Context Modeling

Open Source 14.0B ↓ 61
🤖
Swallow-7b-NVE-instruct-hf
tokyotech-llm

Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.

Open Source 7.0B ↓ 60
🤖
Llama-3.1-8B-Dragonfly-v2
togethercomputer

Note: Users are permitted to use this model in accordance with the Llama 3.1 Community License Agreement.

Open Source 8.0B ↓ 60
🤖
Qianfan-VL-8B
baidu

Qianfan-VL: Domain-Enhanced Universal Vision-Language Models

Multimodal 8.0B ↓ 60
🤖
bilingual-gpt-neox-4b-8k
rinna

Notice: This model requires transformers =4.31.0 to work properly.

Open Source 4.0B ↓ 59
🤖
youri-7b-instruction
rinna

Overview The model is the instruction-tuned version of rinna/youri-7b . It adopts the Alpaca input format.

Open Source 7.0B ↓ 58
🤖
Swallow-70b-hf
tokyotech-llm

Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.

Open Source 70.0B ↓ 58
🤖
Aurora-Spec-Minimax-M2.1
togethercomputer

This is an EAGLE3 draft model trained from scratch (random initialization) using the Aurora inference-time training framework for speculative decoding. Unlike traditional approaches that fine-tune pre-trained models, this model is built entirely through Aurora's online training p…

Open Source ↓ 58