AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

377 models for "MoE" Compare
🤖
GLM-4.5V
zai-org

This model is part of the GLM-V family of models, introduced in the paper GLM-4.1V-Thinking and GLM-4.5V: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

Open Source ↓ 63.2K
🤖
Moonlight-16B-A3B-Instruct
moonshotai

Tech Report HuggingFace Megatron(coming soon)

Open Source 16.0B ↓ 62.2K
🤖
granite-3.0-1b-a400m-instruct
ibm-granite

Model Summary: Granite-3.0-1B-A400M-Instruct is an 1B parameter model finetuned from Granite-3.0-1B-A400M-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets. This model is developed using a diverse set…

Open Source 1.3B ↓ 55.3K
🤖
InternVL3_5-4B
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 4.7B ↓ 55.1K
🤖
ERNIE-4.5-VL-28B-A3B-PT
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Multimodal 28.0B ↓ 54.6K
🤖
Hunyuan-A13B-Instruct
tencent

🤗  Hugging Face       🖥️  Official Website       🕖  HunyuanAPI       🕹️  Demo       🤖  ModelScope

Open Source 13.0B ↓ 48.2K
🤖
InternVL3_5-8B
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 8.5B ↓ 47.7K
🤖
Kimi-VL-A3B-Thinking
moonshotai

[!Warning] This model has a new version: Kimi-VL-A3B-Thinking-2506. Please consider using the new 2506 version for better abilties on general visual understanding, reasoning, video and agent scenarios.

Multimodal 16.0B ↓ 45.9K
🤖
granite-3.1-8b-instruct
ibm-granite

Model Summary: Granite-3.1-8B-Instruct is a 8B parameter long-context instruct model finetuned from Granite-3.1-8B-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets tailored for solving long context pr…

Open Source 8.0B ↓ 45.3K
🤖
command-a-plus-05-2026-bf16
CohereLabs

Command A+ is an open source model with 25 billion active parameters and 218B total parameters model optimized for agentic, multilingual, and reasoning-heavy tasks with a focus on enterprise performance, while also providing support for vision inputs for processing image inputs.

Open Source 218.0B ↓ 39.9K
🤖
academic-ds-9B
ByteDance-Seed

This is a 9B model whose architecture is deepseek-v3, trained from scratch using 350B+ tokens from fully open-source, English-only datasets. It is designed for development and debugging purposes within the open-source community.

Open Source 9.0B ↓ 39.3K
🤖
Kimi-K2-Instruct-0905
moonshotai

📰   Tech Blog         📄   Paper

Open Source ↓ 38.4K