AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

63 models for "Open Weights" Compare
NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 logo
NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
nvidia

:--- :--- Total Parameters 550B (55B active) Architecture LatentMoE - Mamba-2 + MoE + Attention hybrid with Multi-Token Prediction (MTP) Context Length Up to 1M tokens Minimum GPU Requirement 8x GB200/B200/GB300/B300, 16x H100, 8x H200 Supported Languages English, French, Spanish…

Open Source 550.0B ★ 23.0 ↓ 683.5K
NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4 logo
NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4
nvidia

:--- :--- Total Parameters 550B (55B active) Architecture LatentMoE - Mamba-2 + MoE + Attention hybrid with Multi-Token Prediction (MTP) Context Length Up to 1M tokens Minimum GPU Requirement 4xGB200, 4xB200, 4x GB300, 4x B300, 8xH100 Supported Languages English, French, Spanish,…

Open Source 550.0B ★ 23.0 ↓ 159.4K
inclusionAI: Ling 3.0 Flash Sante logo
inclusionAI: Ling 3.0 Flash Sante
inclusionai

Ling 3.0 Flash Sante is a health and medicine-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B total. It is designed for...

Closed Source 124.0B ★ 20.0
Qwen: Qwen3.5-122B-A10B logo
Qwen: Qwen3.5-122B-A10B
qwen

The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. In terms of...

Multimodal 122.0B ★ 18.0
NVIDIA-Nemotron-3-Super-120B-A12B-BF16 logo
NVIDIA-Nemotron-3-Super-120B-A12B-BF16
nvidia

:--- :--- Total Parameters 120B (12B active) Architecture LatentMoE - Mamba-2 + MoE + Attention hybrid with Multi-Token Prediction (MTP) Context Length Up to 1M tokens Minimum GPU Requirement 8× H100-80GB Supported Languages English, French, German, Italian, Japanese, Spanish, Ch…

Reasoning 120.0B ★ 13.0 ↓ 1M
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 logo
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4
nvidia

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4

Open Source 30.0B ★ 13.0 ↓ 839.9K
NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 logo
NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4
nvidia

:--- :--- Total Parameters 120B (12B active) Architecture LatentMoE - Mamba-2 + MoE + Attention hybrid with Multi-Token Prediction (MTP) Context Length Up to 1M tokens Minimum GPU Requirement 1× B200 OR 1× DGX Spark Supported Languages English, French, German, Italian, Japanese,…

Reasoning 120.0B ★ 13.0 ↓ 678.6K
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 logo
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
nvidia

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Open Source 30.0B ★ 13.0 ↓ 599K
NVIDIA-Nemotron-3-Super-120B-A12B-FP8 logo
NVIDIA-Nemotron-3-Super-120B-A12B-FP8
nvidia

:--- :--- Total Parameters 120B (12B active) Architecture LatentMoE - Mamba-2 + MoE + Attention hybrid with Multi-Token Prediction (MTP) Context Length Up to 1M tokens Minimum GPU Requirement 2× H100-80GB Supported Languages English, French, German, Italian, Japanese, Spanish, Ch…

Reasoning 120.0B ★ 13.0 ↓ 64K
MiniCPM5-2B logo
MiniCPM5-2B
openbmb

MiniCPM Tech Report MiniCPM Wiki(Chinese) GitHub Repo UltraData Online Demo

Reasoning 2.52B ★ 12.0 ↓ 1.3M
North-Mini-Code-1.0 logo
North-Mini-Code-1.0
CohereLabs

North Mini Code is an open weights research release of a 30B-A3B parameter model optimized for code generation, agentic software engineering, and terminal tasks.

Code ★ 10.0 ↓ 9.2K
North-Mini-Code-1.0-fp8 logo
North-Mini-Code-1.0-fp8
CohereLabs

North Mini Code is an open weights research release of a 30B-A3B parameter model optimized for code generation, agentic software engineering, and terminal tasks.

Code 30.0B ★ 10.0 ↓ 1.1K