AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,294 models for "Transformer" Compare
🤖
AHN-DN-for-Qwen-2.5-Instruct-3B
ByteDance-Seed

AHN: Artificial Hippocampus Networks for Efficient Long-Context Modeling

Open Source 3.0B ↓ 48
🤖
Qianfan-VL-70B
baidu

Qianfan-VL: Domain-Enhanced Universal Vision-Language Models

Multimodal 70.0B ↓ 47
🤖
ERNIE-4.5-21B-A3B-Base-Paddle
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Open Source 21.0B ↓ 47
🤖
diafill-llm-jp-3.1-13b-instruct4
sbintuitions

Model Summary DiaFill is a Japanese dialogue script generation model designed to produce natural, spoken-style dialogue scripts rich in fillers and brief utternaces. Unlike typical assistant models that respond to users, this model is fine-tuned to generate a multi-turn dialogue…

Open Source 13.0B ↓ 46
🤖
ERNIE-4.5-VL-28B-A3B-Base-Paddle
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Multimodal 28.0B ↓ 46
🤖
ERNIE-4.5-VL-28B-A3B-Paddle
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Multimodal 28.0B ↓ 45
🤖
LongCat-Flash-Omni-FP8
meituan-longcat

Model Introduction We introduce LongCat-Flash-Omni , a state-of-the-art open-source omni-modal model with 560 billion parameters (with 27B activated), excelling at real-time audio-visual interaction, which is attained by leveraging LongCat-Flash's high-performance Shortcut-connec…

Open Source ↓ 44
🤖
EXAONE-Deep-2.4B-AWQ
LGAI-EXAONE

We introduce EXAONE Deep, which exhibits superior capabilities in various reasoning tasks including math and coding benchmarks, ranging from 2.4B to 32B parameters developed and released by LG AI Research. Evaluation results show that 1) EXAONE Deep 2.4B outperforms other models…

Open Source 2.4B ↓ 40
🤖
Klear-46B-A2.5B-Base
Kwai-Klear

🤗 Hugging Face 💻 Github Repository 📑 Technique Report 💬 Issues & Discussions

Open Source 46.0B ↓ 40
🤖
Qwen-SEA-Guard-8B-2602
aisingapore

SEA-Safeguard is a collection of safety-focused Large Language Models (LLMs) built upon the SEA-LION family, designed specifically for the Southeast Asia (SEA) region.

Open Source 8.0B ↓ 38
🤖
distill-bloom-1b3-10x
bigscience

WARNING: This is an intermediary checkpoint and WIP project. It is not fully trained yet. You might want to use Bloom-1B3 if you want a model that has completed training. This model is a distilled version of Bloom-1B3 (10x distillation)

Open Source ↓ 37
🤖
SEA-LION-v1-7B-IT-GPTQ
aisingapore

SEA-LION is a collection of Large Language Models (LLMs) which has been pretrained and instruct-tuned for the Southeast Asia (SEA) region. The sizes of the models range from 3 billion to 7 billion parameters.

Open Source 7.0B ↓ 33