AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

377 models for "MoE" Compare
🤖
Qwen3-VL-235B-A22B-Instruct
Qwen

Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.

Multimodal 235.0B ↓ 825.5K
🤖
MiniMax-M2.5-NVFP4
nvidia

Description: The NVIDIA MiniMax-M2.5-NVFP4 model is the quantized version of MiniMax's MiniMax-M2.5 model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA MiniMax-M2.5 NVFP4 model is q…

Open Source ↓ 776.5K
🤖
Qwen3-Coder-30B-A3B-Instruct
Qwen

Qwen3-Coder is available in multiple sizes. Today, we're excited to introduce Qwen3-Coder-30B-A3B-Instruct . This streamlined model maintains impressive performance and efficiency, featuring the following key enhancements:

Code 30.0B ↓ 774.8K
🤖
DeepSeek-Coder-V2-Lite-Instruct
deepseek-ai

DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence

Code ↓ 717.8K
🤖
Phi-3.5-vision-instruct
microsoft

Phi-3.5-vision is a lightweight, state-of-the-art open multimodal model built upon datasets which include - synthetic data and filtered publicly available websites - with a focus on very high-quality, reasoning dense data both on text and vision. The model belongs to the Phi-3 mo…

Multimodal 4.2B ↓ 681K
🤖
Kimi-K2.6
moonshotai

🤗   huggingchat     📰   Tech Blog

Open Source ↓ 647.5K
🤖
Phi-3-mini-4k-instruct
microsoft

🎉 Phi-3.5 : [[mini-instruct]](https://huggingface.co/microsoft/Phi-3.5-mini-instruct); [[MoE-instruct]](https://huggingface.co/microsoft/Phi-3.5-MoE-instruct) ; [[vision-instruct]](https://huggingface.co/microsoft/Phi-3.5-vision-instruct)

Open Source ↓ 612.3K
🤖
DeepSeek-R1-Distill-Qwen-32B
deepseek-ai

We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With R…

Reasoning 32.0B ↓ 577.8K
🤖
Kimi-K2.5
moonshotai

📰   Tech Blog         📄   Paper

Open Source ↓ 554.9K
🤖
DeepSeek-R1-Distill-Qwen-1.5B
deepseek-ai

We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With R…

Reasoning 1.5B ↓ 488.7K
🤖
MiniMax-M2.5
MiniMaxAI

Join Our 💬 WeChat 🧩 Discord community. MiniMax Agent ⚡️ API MCP MiniMax Website 🤗 Hugging Face 🚀 Hugging Face API 🐙 GitHub 🤖️ ModelScope 📄 License: Modified-MIT

Open Source ↓ 487.4K
🤖
DeepSeek-R1-Distill-Qwen-14B
deepseek-ai

We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With R…

Reasoning 14.0B ↓ 460.1K