AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

377 models for "MoE" Compare
🤖
granite-swash-3b-a600m
ibm-granite

Granite-SWASH-3B-a600M (Sliding Window Attention + Sinks Hybrid)

Open Source 3.0B ↓ 6.3K
🤖
granite-3.1-3b-a800m-instruct
ibm-granite

Model Summary: Granite-3.1-3B-A800M-Instruct is a 3B parameter long-context instruct model finetuned from Granite-3.1-3B-A800M-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets tailored for solving lon…

Open Source 3.0B ↓ 5.8K
🤖
InternVL3_5-8B-Instruct
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 8.5B ↓ 5.6K
🤖
MiniMax-M1-40k
MiniMaxAI

We introduce MiniMax-M1, the world's first open-weight, large-scale hybrid-attention reasoning model. MiniMax-M1 is powered by a hybrid Mixture-of-Experts (MoE) architecture combined with a lightning attention mechanism. The model is developed based on our previous MiniMax-Text-0…

Open Source ↓ 5.5K
🤖
Flex-reddit-2x7B-1T
allenai

FlexOlmo is a new kind of LM that unlocks a new paradigm of data collaboration. With FlexOlmo, data owners can contribute to the development of open language models without giving up control of their data. There is no need to share raw data directly, and data contributors can dec…

Open Source 7.0B ↓ 5.4K
🤖
GLM-4.6V-FP8
zai-org

This model is part of the GLM-V family of models, introduced in the paper GLM-4.1V-Thinking and GLM-4.5V: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

Open Source ↓ 5.3K
🤖
deepseek-vl2
deepseek-ai

Introducing DeepSeek-VL2, an advanced series of large Mixture-of-Experts (MoE) Vision-Language Models that significantly improves upon its predecessor, DeepSeek-VL. DeepSeek-VL2 demonstrates superior capabilities across various tasks, including but not limited to visual question…

Multimodal 27.0B ↓ 5.2K
🤖
granite-3.1-1b-a400m-instruct
ibm-granite

Model Summary: Granite-3.1-1B-A400M-Instruct is a 1B parameter long-context instruct model finetuned from Granite-3.1-1B-A400M-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets tailored for solving lon…

Open Source 1.0B ↓ 5.2K
🤖
Baichuan-M3-235B-GPTQ-INT4
baichuan-inc

From Inquiry to Decision: Building Trustworthy Medical AI

Open Source 235.0B ↓ 5.1K
🤖
Jamba-tiny-random
ai21labs

This is a tiny, dummy version of Jamba, used for debugging and experimentation over the Jamba architecture.

Open Source 0.128B ↓ 5K
🤖
GLM-4-32B-0414
zai-org

The GLM family welcomes new members, the GLM-4-32B-0414 series models, featuring 32 billion parameters. Its performance is comparable to OpenAI’s GPT series and DeepSeek’s V3/R1 series. It also supports very user-friendly local deployment features. GLM-4-32B-Base-0414 was pre-tra…

Open Source 32.0B ↓ 4.4K
🤖
command-a-plus-05-2026-w4a4
CohereLabs

Command A+ is an open source model with 25 billion active parameters and 218B total parameters model optimized for agentic, multilingual, and reasoning-heavy tasks with a focus on enterprise performance, while also providing support for vision inputs for processing image inputs.

Open Source 218.0B ↓ 4.4K