AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

377 models for "MoE" Compare
🤖
NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-NVFP4
nvidia

NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-NVFP4

Open Source 75.0B ↓ 437.2K
🤖
GLM-5.1-FP8
zai-org

👋 Join our WeChat or Discord community. 📖 Check out the GLM-5.1 blog and GLM-5 Technical report . 📍 Use GLM-5.1 API services on Z.ai API Platform. 🔜 GLM-5.1 will be available on chat.z.ai in the coming days.

Open Source ↓ 398K
🤖
DeepSeek-V2-Lite-Chat
deepseek-ai

Model Download Evaluation Results Model Architecture API Platform License Citation

Open Source ↓ 396.7K
🤖
DeepSeek-R1-Distill-Llama-8B
deepseek-ai

We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With R…

Reasoning 8.0B ↓ 388.6K
🤖
DeepSeek-R1-Distill-Qwen-7B
deepseek-ai

We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With R…

Reasoning 7.0B ↓ 374K
🤖
DeepSeek-V2-Lite
deepseek-ai

Model Download Evaluation Results Model Architecture API Platform License Citation

Open Source ↓ 368.6K
🤖
Kimi-VL-A3B-Instruct
moonshotai

📄 Tech Report     📄 Github     💬 Chat Web

Multimodal 16.0B ↓ 315.8K
🤖
Phi-3.5-mini-instruct
microsoft

🎉 Phi-4 : [multimodal-instruct onnx]; [mini-instruct onnx]

Open Source ↓ 300.7K
🤖
Phi-tiny-MoE-instruct
microsoft

Phi-tiny-MoE is a lightweight Mixture of Experts (MoE) model with 3.8B total parameters and 1.1B activated parameters. It is compressed and distilled from the base model shared by Phi-3.5-MoE and GRIN-MoE using the SlimMoE approach, then post-trained via supervised fine-tuning an…

Open Source ↓ 272.7K
🤖
deepseek-vl2-tiny
deepseek-ai

Introducing DeepSeek-VL2, an advanced series of large Mixture-of-Experts (MoE) Vision-Language Models that significantly improves upon its predecessor, DeepSeek-VL. DeepSeek-VL2 demonstrates superior capabilities across various tasks, including but not limited to visual question…

Multimodal 3.37B ↓ 265.9K
🤖
OLMoE-1B-7B-0125-Instruct
allenai

OLMoE-1B-7B-0125-Instruct January 2025 is post-trained variant of the OLMoE-1B-7B January 2025 model, which has undergone supervised finetuning on an OLMo-specific variant of the Tülu 3 dataset and further DPO training on this dataset, and finally RLVR training using this data. T…

Open Source 1.0B ↓ 264.8K
🤖
MiniMax-M2
MiniMaxAI

Join Our 💬 WeChat 🧩 Discord community. MiniMax Agent ⚡️ API (Now Free for a limited time!) MCP MiniMax Website 🤗 Hugging Face 🐙 GitHub 🤖️ ModelScope 📄 License: MIT

Open Source ↓ 258.3K