AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

377 models for "MoE" Compare
🤖
InternVL3_5-241B-A28B-Instruct
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 241.0B ↓ 1K
🤖
GLM-4-32B-Base-0414
zai-org

The GLM family welcomes new members, the GLM-4-32B-0414 series models, featuring 32 billion parameters. Its performance is comparable to OpenAI’s GPT series and DeepSeek’s V3/R1 series. It also supports very user-friendly local deployment features. GLM-4-32B-Base-0414 was pre-tra…

Open Source 32.0B ↓ 1K
🤖
nanowhale-100m
HuggingFaceTB

A small ~110M parameter language model implementing the DeepSeek-V4 architecture , fine-tuned for chat/instruction following. Trained from scratch — no weights from DeepSeek-V4 were used.

Open Source ↓ 939
🤖
InternVL3_5-1B-Instruct
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 1.0B ↓ 882
🤖
LLaDA2.0-mini-CAP
inclusionAI

LLaDA2.0-mini-CAP is an enhanced version of LLaDA2.0-mini that incorporates Confidence-Aware Parallel (CAP) Training for significantly improved inference efficiency. Built upon the 16B-A1B Mixture-of-Experts (MoE) diffusion architecture, this model achieves faster parallel decodi…

Open Source ↓ 832
🤖
InternVL3_5-8B-Pretrained
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 8.0B ↓ 735
🤖
GLM-4.5-Base
zai-org

👋 Join our Discord community. 📖 Check out the GLM-4.5 technical blog , technical report , and Zhipu AI technical documentation . 📍 Use GLM-4.5 API services on Z.ai API Platform (Global) or Zhipu AI Open Platform (Mainland China) . 👉 One click to GLM-4.5 .

Open Source ↓ 729
🤖
nanowhale-100m-base
HuggingFaceTB

A small ~110M parameter language model implementing the DeepSeek-V4 architecture from scratch. This is the pretrained base model — see HuggingFaceTB/nanowhale-100m for the SFT/chat version.

Open Source ↓ 706
🤖
LLaDA2.0-mini-preview
inclusionAI

LLaDA2.0-mini-preview is a diffusion language model featuring a 16BA1B Mixture-of-Experts (MoE) architecture. As an enhanced, instruction-tuned iteration of the LLaDA series, it is optimized for practical applications.

Open Source ↓ 618
🤖
Hunyuan-1.8B-Instruct
tencent

🤗  HuggingFace     🤖  ModelScope     🪡  AngelSlim

Open Source 1.8B ↓ 575
🤖
LLaDA2.1-flash
inclusionAI

🚀 LLaDA2.1-flash is now live on ZenmuxAI ! Try it via API 🛠️ or Chat 💬: https://zenmux.ai/inclusionai/llada2.1-flash

Open Source ↓ 548
🤖
Ling-2.6-flash-base
inclusionAI

🤗 Hugging Face      🤖 ModelScope       Tech Report       💻 GitHub

Open Source ↓ 545