AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

377 models for "MoE" Compare
🤖
InternVL3_5-38B-Flash
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 38.0B ↓ 527
🤖
InternVL3_5-8B-Flash
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 8.0B ↓ 508
🤖
DeepSeek-V3.1-Terminus-W4AFP8
tencent

This model is a mixed-precision quantized version of DeepSeek-V3.1-Terminus, with dense layer keep the FP8 quantization of the original model, while MoE layers uses INT4 weights and FP8 activation, also called W4AFP8.

Open Source ↓ 483
🤖
ERNIE-4.5-VL-424B-A47B-PT
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Multimodal 424.0B ↓ 449
🤖
Hy-Embodied-VLM-1.0
tencent

Hy-Embodied-VLM-1.0 Efficient Physical-World Agents Tencent Robotics X × Hy Vision Team × Futian Laboratory

Multimodal ↓ 449
🤖
MiMo-V2.5-Pro-Base
XiaomiMiMo

🤗 HuggingFace   📰 Blog   🎨 Xiaomi MiMo API Platform   🗨️ Xiaomi MiMo Studio  

Open Source ↓ 430
🤖
LLaDA-MoE-7B-A1B-Base
inclusionAI

This model is based on the principles described in the paper Large Language Diffusion Models.

Open Source 7.0B ↓ 411
🤖
MiMo-Embodied-7B
XiaomiMiMo

🤗 HuggingFace   📔 Technical Report  

Open Source 7.0B ↓ 409
🤖
Tencent-Hunyuan-Large
tencent

&nbsp GITHUB &nbsp&nbsp &nbsp&nbsp🖥️&nbsp&nbsp official website &nbsp&nbsp|&nbsp&nbsp🕖&nbsp&nbsp HunyuanAPI |&nbsp&nbsp🐳&nbsp&nbsp Gitee Technical Report &nbsp&nbsp|&nbsp&nbsp Demo &nbsp&nbsp&nbsp|&nbsp&nbsp Tencent Cloud TI &nbsp&nbsp&nbsp

Open Source ↓ 405
🤖
Stable-DiffCoder-8B-Instruct
ByteDance-Seed

Introduction We are thrilled to introduce Stable-DiffCoder, which is a strong code diffusion large language model. Built directly on the Seed-Coder architecture, data, and training pipeline, it introduces a block diffusion continual pretraining (CPT) stage with a tailored warmup…

Code 8.0B ↓ 398
🤖
Hunyuan-0.5B-Instruct
tencent

🤗  HuggingFace     🤖  ModelScope     🪡  AngelSlim

Open Source 0.5B ↓ 392
🤖
Hunyuan-4B-Instruct
tencent

🤗  HuggingFace     🤖  ModelScope     🪡  AngelSlim

Open Source 4.0B ↓ 338