AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,877 models Compare
🤖
Llama-3.1-8B-Dragonfly-Med-v2
togethercomputer

Note: Users are permitted to use this model in accordance with the Llama 3.1 Community License Agreement. Additionally, due to the licensing restrictions of the dataset used to train this model, which prohibits commercial use, the Dragonfly-Med model is restricted to non-commerci…

Open Source 8.0B ↓ 61
🤖
AHN-GDN-for-Qwen-2.5-Instruct-14B
ByteDance-Seed

AHN: Artificial Hippocampus Networks for Efficient Long-Context Modeling

Open Source 14.0B ↓ 61
🤖
Swallow-7b-NVE-instruct-hf
tokyotech-llm

Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.

Open Source 7.0B ↓ 60
🤖
Llama-3.1-8B-Dragonfly-v2
togethercomputer

Note: Users are permitted to use this model in accordance with the Llama 3.1 Community License Agreement.

Open Source 8.0B ↓ 60
🤖
Qianfan-VL-8B
baidu

Qianfan-VL: Domain-Enhanced Universal Vision-Language Models

Multimodal 8.0B ↓ 60
🤖
bilingual-gpt-neox-4b-8k
rinna

Notice: This model requires transformers =4.31.0 to work properly.

Open Source 4.0B ↓ 59
🤖
youri-7b-instruction
rinna

Overview The model is the instruction-tuned version of rinna/youri-7b . It adopts the Alpaca input format.

Open Source 7.0B ↓ 58
🤖
Swallow-70b-hf
tokyotech-llm

Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.

Open Source 70.0B ↓ 58
🤖
Aurora-Spec-Minimax-M2.1
togethercomputer

This is an EAGLE3 draft model trained from scratch (random initialization) using the Aurora inference-time training framework for speculative decoding. Unlike traditional approaches that fine-tune pre-trained models, this model is built entirely through Aurora's online training p…

Open Source ↓ 58
🤖
mt0-xxl-mt
bigscience

1. Model Summary 2. Use 3. Limitations 4. Training 5. Evaluation 7. Citation

Open Source ↓ 58
🤖
ERNIE-4.5-0.3B-Base-Paddle
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Open Source 0.3B ↓ 58
🤖
CodeLlama-70b-hf
meta-llama

Code 70.0B ↓ 58