AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

377 models for "MoE" Compare
🤖
LongCat-Flash-Thinking-2601-FP8
meituan-longcat

We introduce an updated version of LongCat-Flash-Thinking-2601, a powerful and efficient Large Reasoning Model (LRM) with 560 billion total parameters, built upon an innovative Mixture-of-Experts (MoE) architecture. Beyond inheriting the domain-parallel training recipe in our pre…

Open Source ↓ 55
🤖
ERNIE-4.5-VL-28B-A3B-Base-PT
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Multimodal 28.0B ↓ 54
🤖
ERNIE-4.5-21B-A3B-Paddle
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Open Source 21.0B ↓ 50
🤖
ERNIE-4.5-21B-A3B-Base-Paddle
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Open Source 21.0B ↓ 47
🤖
ERNIE-4.5-VL-28B-A3B-Base-Paddle
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Multimodal 28.0B ↓ 46
🤖
ERNIE-4.5-VL-28B-A3B-Paddle
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Multimodal 28.0B ↓ 45
🤖
LongCat-Flash-Omni-FP8
meituan-longcat

Model Introduction We introduce LongCat-Flash-Omni , a state-of-the-art open-source omni-modal model with 560 billion parameters (with 27B activated), excelling at real-time audio-visual interaction, which is attained by leveraging LongCat-Flash's high-performance Shortcut-connec…

Open Source ↓ 44
🤖
LongCat-Flash-Thinking-FP8
meituan-longcat

We introduce and release LongCat-Flash-Thinking , which is a powerful and efficient large reasoning model (LRM) with 560 billion total parameters, featuring an innovative Mixture-of-Experts (MoE) architecture. The model incorporates a dynamic computation mechanism that activates…

Open Source ↓ 41
🤖
Klear-46B-A2.5B-Base
Kwai-Klear

🤗 Hugging Face 💻 Github Repository 📑 Technique Report 💬 Issues & Discussions

Open Source 46.0B ↓ 40
🤖
Klear-46B-A2.5B-Instruct
Kwai-Klear

🤗 Hugging Face 💻 Github Repository 📑 Technique Report 💬 Issues & Discussions

Open Source 46.0B ↓ 33
🤖
ERNIE-4.5-VL-424B-A47B-Paddle
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Multimodal 424.0B ↓ 32
🤖
Qwen: Qwen3 Coder 30B A3B Instruct
qwen

Qwen3-Coder-30B-A3B-Instruct is a 30.5B parameter Mixture-of-Experts (MoE) model with 128 experts (8 active per forward pass), designed for advanced code generation, repository-scale understanding, and agentic tool use. Built on the...

Code