AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

377 models for "MoE" Compare
🤖
granite-3.0-1b-a400m-base
ibm-granite

Model Summary: Granite-3.0-1B-A400M-Base is a decoder-only language model to support a variety of text-to-text generation tasks. It is trained from scratch following a two-stage training strategy. In the first stage, it is trained on 8 trillion tokens sourced from diverse domains…

Open Source 1.3B ↓ 37.6K
🤖
DeepSeek-V2
deepseek-ai

Model Download Evaluation Results Model Architecture API Platform License Citation

Open Source ↓ 37K
🤖
deepseek-moe-16b-chat
deepseek-ai

[🏠Homepage] [🤖 Chat with DeepSeek LLM] [Discord] [Wechat(微信)]

Open Source 16.4B ↓ 34.4K
🤖
LongCat-Flash-Chat
meituan-longcat

Model Introduction We introduce LongCat-Flash, a powerful and efficient language model with 560 billion total parameters, featuring an innovative Mixture-of-Experts (MoE) architecture. The model incorporates a dynamic computation mechanism that activates 18.6B∼31.3B parameters (a…

Open Source ↓ 34K
🤖
OLMo-2-1124-7B-SFT
allenai

Upon the initial release of OLMo-2 models, we realized the post-trained models did not share the pre-tokenization logic that the base models use. As a result, we have trained new post-trained models. The new models are available under the same names as the original models, but we…

Open Source 7.0B ↓ 33.8K
🤖
GLM-4-9B-0414
zai-org

The GLM family welcomes new members, the GLM-4-32B-0414 series models, featuring 32 billion parameters. Its performance is comparable to OpenAI’s GPT series and DeepSeek’s V3/R1 series. It also supports very user-friendly local deployment features. GLM-4-32B-Base-0414 was pre-tra…

Open Source 9.0B ↓ 32.7K
🤖
granite-3.1-2b-instruct
ibm-granite

Model Summary: Granite-3.1-2B-Instruct is a 2B parameter long-context instruct model finetuned from Granite-3.1-2B-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets tailored for solving long context pr…

Open Source 2.0B ↓ 31K
🤖
InternVL3_5-30B-A3B
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 30.85B ↓ 28.5K
🤖
ERNIE-4.5-21B-A3B-PT
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Open Source 21.0B ↓ 27.4K
🤖
InternVL3_5-38B
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 38.4B ↓ 24.1K
🤖
granite-4.0-h-micro
ibm-granite

📣 Update [10-07-2025]: Added a default system prompt to the chat template to guide the model towards more professional, accurate, and safe responses.

Open Source 3.0B ↓ 20.1K
🤖
Nous-Hermes-2-Mixtral-8x7B-DPO
NousResearch

Nous Hermes 2 Mixtral 8x7B DPO is the new flagship Nous Research model trained over the Mixtral 8x7B MoE LLM.

Open Source 7.0B ↓ 19.6K