AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

375 models for "MoE" Compare
🤖
NVIDIA-Nemotron-3-Super-120B-A12B-BF16
nvidia

:--- :--- Total Parameters 120B (12B active) Architecture LatentMoE - Mamba-2 + MoE + Attention hybrid with Multi-Token Prediction (MTP) Context Length Up to 1M tokens Minimum GPU Requirement 8× H100-80GB Supported Languages English, French, German, Italian, Japanese, Spanish, Ch…

Reasoning 120.0B ★ 26.0 ↓ 1.1M
🤖
gemma-4-26B-A4B-it-qat-q4_0-unquantized
google

Hugging Face GitHub Launch Blog Documentation Technical Report License : Apache 2.0 Authors : Google DeepMind

Open Source 26.0B ★ 26.0 ↓ 239.8K
🤖
NVIDIA-Nemotron-3-Super-120B-A12B-FP8
nvidia

:--- :--- Total Parameters 120B (12B active) Architecture LatentMoE - Mamba-2 + MoE + Attention hybrid with Multi-Token Prediction (MTP) Context Length Up to 1M tokens Minimum GPU Requirement 2× H100-80GB Supported Languages English, French, German, Italian, Japanese, Spanish, Ch…

Reasoning 120.0B ★ 26.0 ↓ 125.6K
🤖
NVIDIA: Nemotron 3 Super (free)
nvidia

NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer...

Closed Source ★ 26.0
🤖
NVIDIA: Nemotron 3 Super
nvidia

NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer...

Closed Source ★ 26.0
🤖
Google: Gemma 4 26B A4B (free)
google

Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...

Closed Source ★ 26.0
🤖
MiMo-V2-Flash
XiaomiMiMo

🤗 HuggingFace   📔 Technical Report   📰 Blog   Play around!   🗨️ Xiaomi MiMo Studio   🎨 Xiaomi MiMo API Platform

Open Source ★ 25.0 ↓ 88K
🤖
Ling-3.0-tiny
inclusionAI

🤗 Hugging Face      🤖 ModelScope      🐙 OpenRouter   

Reasoning 7.9B ★ 25.0 ↓ 21.6K
🤖
MiMo-V2-Flash-Base
XiaomiMiMo

🤗 HuggingFace   📔 Technical Report   📰 Blog   Play around!   🗨️ Xiaomi MiMo Studio   🎨 Xiaomi MiMo API Platform

Open Source ★ 25.0 ↓ 354
🤖
gpt-oss-120b
openai

Try gpt-oss · Guides · Model card · OpenAI blog

Open Source 120.0B ★ 24.0 ↓ 5.4M
🤖
Qwen3.5-35B-A3B
Qwen

[!Note] This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.

Open Source 35.0B ★ 24.0 ↓ 2.5M
🤖
Qwen3.5-35B-A3B-FP8
Qwen

[!Note] This repository contains FP8-quantized model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc. The quantization method is fin…

Multimodal 35.0B ★ 24.0 ↓ 1.4M