AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

377 models for "MoE" Compare
🤖
Molmo-72B-0924
allenai

Molmo is a family of open vision-language models developed by the Allen Institute for AI. Molmo models are trained on PixMo, a dataset of 1 million, highly-curated image-text pairs. It has state-of-the-art performance among multimodal models with a similar size while being fully…

Multimodal 72.0B ↓ 4.2K
🤖
MiniCPM-MoE-8x2B
openbmb

The MiniCPM-MoE-8x2B is a decoder-only transformer-based generative language model.

Open Source 13.6B ↓ 4.2K
🤖
Ring-mini-linear-2.0
inclusionAI

📖 Technical Report &nbsp&nbsp &nbsp&nbsp 🤗 Hugging Face &nbsp&nbsp &nbsp&nbsp🤖 ModelScope

Open Source 16.4B ↓ 4K
🤖
LLaDA2.0-flash
inclusionAI

LLaDA2.0-flash is a diffusion language model featuring a 100BA6B Mixture-of-Experts (MoE) architecture. As an enhanced, instruction-tuned iteration of the LLaDA2.0 series, it is optimized for practical applications.

Open Source ↓ 3.7K
🤖
granite-3.1-1b-a400m-base
ibm-granite

Model Summary: Granite-3.1-1B-A400M-Base extends the context length of Granite-3.0-1B-A400M-Base from 4K to 128K using a progressive training strategy by increasing the supported context length in increments while adjusting RoPE theta until the model has successfully adapted to d…

Open Source 1.0B ↓ 3.7K
🤖
OLMo-2-1124-7B-DPO
allenai

Upon the initial release of OLMo-2 models, we realized the post-trained models did not share the pre-tokenization logic that the base models use. As a result, we have trained new post-trained models. The new models are available under the same names as the original models, but we…

Open Source 7.0B ↓ 3.5K
🤖
InternVL3_5-38B-Instruct
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 38.0B ↓ 3.2K
🤖
InternVL3_5-241B-A28B
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 241.0B ↓ 3.2K
🤖
LongCat-Flash-Thinking-2601
meituan-longcat

We introduce an updated version of LongCat-Flash-Thinking, a powerful and efficient Large Reasoning Model (LRM) with 560 billion total parameters, built upon an innovative Mixture-of-Experts (MoE) architecture. Beyond inheriting the domain-parallel training recipe in our previous…

Open Source ↓ 3K
🤖
DeepSeek-Coder-V2-Base
deepseek-ai

DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence

Code ↓ 2.9K
🤖
MiniMax-Text-01
MiniMaxAI

MiniMax-Text-01 is a powerful language model with 456 billion total parameters, of which 45.9 billion are activated per token. To better unlock the long context capabilities of the model, MiniMax-Text-01 adopts a hybrid architecture that combines Lightning Attention, Softmax Atte…

Open Source ↓ 2.8K
🤖
Nous-Hermes-2-Mixtral-8x7B-SFT
NousResearch

Nous Hermes 2 Mixtral 8x7B SFT is the supervised finetune only version of our new flagship Nous Research model trained over the Mixtral 8x7B MoE LLM.

Open Source 7.0B ↓ 2.8K