AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

375 models for "MoE" Compare
🤖
DeepSeek-V4-Flash-DSpark
deepseek-ai

DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence

Open Source ★ 52.0 ↓ 381.7K
🤖
DeepSeek-V4-Flash-NVFP4
nvidia

Description: The NVIDIA DeepSeek-V4-Flash-NVFP4 model is a quantized version of DeepSeek AI's DeepSeek-V4-Flash model, an autoregressive Mixture-of-Experts language model that uses an optimized Transformer architecture with hybrid attention (Compressed Sparse Attention and Heavil…

Open Source ★ 52.0 ↓ 253.4K
🤖
Kimi-K2.7-Code-NVFP4
nvidia

Description: The NVIDIA Kimi-K2.7-Code NVFP4 model is the quantized version of the Moonshot AI's Kimi-K2.7-Code model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Kimi-K2.7-Code NV…

Code ★ 43.0 ↓ 489.3K
🤖
Kimi-K2.7-Code
moonshotai

Kimi K2.7 Code is a coding-focused agentic model built upon Kimi K2.6. With substantial improvements on real-world long-horizon coding tasks, it strengthens end-to-end task completion across complex software engineering workflows while improving token efficiency, reducing thinkin…

Code ★ 43.0 ↓ 229.3K
🤖
MiMo-V2.5-Pro
XiaomiMiMo

🤗 HuggingFace   📰 Blog   🎨 Xiaomi MiMo API Platform   🗨️ Xiaomi MiMo Studio  

Open Source ★ 43.0 ↓ 48.2K
🤖
MiMo-V2.5-Pro-FP4-DFlash
XiaomiMiMo

🎨 Xiaomi MiMo API Platform (Request Access)     🗨️ Xiaomi MiMo Studio (Free Trial)

Open Source ★ 43.0 ↓ 491
🤖
MoonshotAI: Kimi K2 0711
moonshotai

Kimi K2 Instruct is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32 billion active per forward pass. It is optimized for...

Closed Source ★ 43.0
🤖
Hy3-preview
tencent

           

Open Source ★ 42.0 ↓ 61.3K
🤖
Hy3-FP8
tencent

           

Open Source ★ 42.0 ↓ 39K
🤖
Hy3
tencent

           

Open Source ★ 42.0 ↓ 11.7K
🤖
NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4
nvidia

:--- :--- Total Parameters 550B (55B active) Architecture LatentMoE - Mamba-2 + MoE + Attention hybrid with Multi-Token Prediction (MTP) Context Length Up to 1M tokens Minimum GPU Requirement 4xGB200, 4xB200, 4x GB300, 4x B300, 8xH100 Supported Languages English, French, Spanish,…

Open Source 550.0B ★ 38.0 ↓ 416K
🤖
MiMo-V2.5
XiaomiMiMo

🤗 HuggingFace   📰 Blog   🎨 Xiaomi MiMo API Platform   🗨️ Xiaomi MiMo Studio  

Open Source ★ 38.0 ↓ 379.1K