AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

377 models for "MoE" Compare
🤖
InternVL3_5-GPT-OSS-20B-A4B-Preview
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 20.0B ★ 15.0 ↓ 30.8K
🤖
Solar-Open-100B
upstage

Solar Open is Upstage's flagship 102B-parameter large language model, trained entirely from scratch and released under the Upstage Solar License (see LICENSE for details). As a Mixture-of-Experts (MoE) architecture, it delivers enterprise-grade performance in reasoning, instructi…

Open Source 100.0B ★ 15.0 ↓ 14.8K
🤖
OpenAI: gpt-oss-20b (batch)
openai

gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for...

Closed Source ★ 15.0
🤖
NVIDIA: Nemotron 3 Nano 30B A3B
nvidia

NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI systems. The model is fully...

Closed Source ★ 15.0
🤖
Llama-4-Maverick-17B-128E-Instruct-FP8
meta-llama

Multimodal 400.0B ★ 14.0 ↓ 75.2K
🤖
Llama-4-Maverick-17B-128E-Instruct
meta-llama

Multimodal 400.0B ★ 14.0 ↓ 11.8K
🤖
Meta: Llama 4 Maverick
meta-llama

Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forward...

Closed Source ★ 14.0
🤖
Upstage: Solar Pro 3
upstage

Solar Pro 3 is Upstage's powerful Mixture-of-Experts (MoE) language model. With 102B total parameters and 12B active parameters per forward pass, it delivers exceptional performance while maintaining computational efficiency. Optimized...

Closed Source ★ 14.0
🤖
diffusiongemma-26B-A4B-it
google

Hugging Face GitHub Launch Blog Documentation License : Apache 2.0 Authors : Google DeepMind

Open Source 26.0B ★ 13.0 ↓ 1.3M
🤖
diffusiongemma-26B-A4B-it-NVFP4
nvidia

--- pipeline tag: text-generation base model: google/diffusiongemma-26B-A4B-it license: apache-2.0 license name: apache-license-2.0 license link: https://ai.google.dev/gemma/apache 2 tags: - nvidia - ModelOpt - DiffusionGemma-26B-A4B-IT - quantized - NVFP4 - nvfp4 ---

Open Source 26.0B ★ 13.0 ↓ 299.1K
🤖
sarvam-105b
sarvamai

Want a smaller model? Download Sarvam-30B!

Open Source 105.0B ★ 12.0 ↓ 6.5K
🤖
sarvam-105b-fp8
sarvamai

Want a smaller model? Download Sarvam-30B!

Open Source 105.0B ★ 12.0 ↓ 1.7K