AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

377 models for "MoE" Compare
🤖
LLaDA2.0-mini
inclusionAI

LLaDA2.0-mini is a diffusion language model featuring a 16BA1B Mixture-of-Experts (MoE) architecture. As an enhanced, instruction-tuned iteration of the LLaDA series, it is optimized for practical applications.

Open Source ↓ 223.3K
🤖
Kimi-K2-Instruct
moonshotai

📰   Tech Blog         📄   Paper

Open Source ↓ 166.4K
🤖
Kimi-K2.6-NVFP4
nvidia

Description: The NVIDIA Kimi-K2.6-NVFP4 model is the quantized version of the Moonshot AI's Kimi-K2.6 model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Kimi-K2.6 NVFP4 model is qu…

Open Source ↓ 162.3K
🤖
Kimi-K2.5-NVFP4
nvidia

Description: The NVIDIA Kimi-K2.5-NVFP4 model is the quantized version of the Moonshot AI's Kimi-K2.5 model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Kimi-K2.5 NVFP4 model is qu…

Open Source ↓ 156.1K
🤖
Step-3.5-Flash
stepfun-ai

Step 3.5 Flash (visit website) is our most capable open-source foundation model, engineered to deliver frontier reasoning and agentic capabilities with exceptional efficiency. Built on a sparse Mixture of Experts (MoE) architecture, it selectively activates only 11B of its 196B p…

Open Source ↓ 155.3K
🤖
granite-3.0-8b-instruct
ibm-granite

Model Summary: Granite-3.0-8B-Instruct is a 8B parameter model finetuned from Granite-3.0-8B-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets. This model is developed using a diverse set of techniques…

Open Source 8.1B ↓ 152.6K
🤖
OLMoE-1B-7B-0924
allenai

OLMoE-1B-7B is a Mixture-of-Experts LLM with 1B active and 7B total parameters released in September 2024 (0924). It yields state-of-the-art performance among models with a similar cost (1B) and is competitive with much larger models like Llama2-13B. OLMoE is 100% open-source.

Open Source 7.0B ↓ 151.2K
🤖
Phi-3.5-MoE-instruct
microsoft

Phi-3.5-MoE is a lightweight, state-of-the-art open model built upon datasets used for Phi-3 - synthetic data and filtered publicly available documents - with a focus on very high-quality, reasoning dense data. The model supports multilingual and comes with 128K context length (i…

Open Source ↓ 134.1K
🤖
GLM-4.5-Air
zai-org

👋 Join our Discord community. 📖 Check out the GLM-4.5 technical blog , technical report , and Zhipu AI technical documentation . 📍 Use GLM-4.5 API services on Z.ai API Platform (Global) or Zhipu AI Open Platform (Mainland China) . 👉 One click to GLM-4.5 .

Open Source ↓ 116.4K
🤖
LLaDA2.1-mini
inclusionAI

🚀 LLaDA2.1-flash is now live on ZenmuxAI ! Try it via API 🛠️ or Chat 💬: https://zenmux.ai/inclusionai/llada2.1-flash

Open Source 16.0B ↓ 108K
🤖
GLM-4.5
zai-org

👋 Join our Discord community. 📖 Check out the GLM-4.5 technical blog , technical report , and Zhipu AI technical documentation . 📍 Use GLM-4.5 API services on Z.ai API Platform (Global) or Zhipu AI Open Platform (Mainland China) . 👉 One click to GLM-4.5 .

Open Source ↓ 95.8K
🤖
DeepSeek-R1-Distill-Llama-70B
deepseek-ai

We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With R…

Reasoning 70.0B ↓ 94K