AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

133 models for "RLHF" Compare
Qwen2.5-7B logo
Qwen2.5-7B
Qwen

Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2:

Open Source 7.61B ★ 23.0 ↓ 786.8K
NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 logo
NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
nvidia

:--- :--- Total Parameters 550B (55B active) Architecture LatentMoE - Mamba-2 + MoE + Attention hybrid with Multi-Token Prediction (MTP) Context Length Up to 1M tokens Minimum GPU Requirement 8x GB200/B200/GB300/B300, 16x H100, 8x H200 Supported Languages English, French, Spanish…

Open Source 550.0B ★ 23.0 ↓ 683.5K
NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4 logo
NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4
nvidia

:--- :--- Total Parameters 550B (55B active) Architecture LatentMoE - Mamba-2 + MoE + Attention hybrid with Multi-Token Prediction (MTP) Context Length Up to 1M tokens Minimum GPU Requirement 4xGB200, 4xB200, 4x GB300, 4x B300, 8xH100 Supported Languages English, French, Spanish,…

Open Source 550.0B ★ 23.0 ↓ 159.4K
granite-4.2-30b logo
granite-4.2-30b
ibm-granite

--- --- Developers Granite Team, IBM Model Type Decoder-only Dense Transformer (Reasoning) Architecture GraniteForCausalLM Base Model Granite-4.1-30B-Base Parameters 30B Context Length Natively Supports 128K (Long-context extension to 512K) Precision bfloat16 Tested Languages Eng…

Open Source 30.0B ★ 15.0 ↓ 38.8K
NVIDIA-Nemotron-3-Super-120B-A12B-BF16 logo
NVIDIA-Nemotron-3-Super-120B-A12B-BF16
nvidia

:--- :--- Total Parameters 120B (12B active) Architecture LatentMoE - Mamba-2 + MoE + Attention hybrid with Multi-Token Prediction (MTP) Context Length Up to 1M tokens Minimum GPU Requirement 8× H100-80GB Supported Languages English, French, German, Italian, Japanese, Spanish, Ch…

Reasoning 120.0B ★ 13.0 ↓ 1M
NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 logo
NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4
nvidia

:--- :--- Total Parameters 120B (12B active) Architecture LatentMoE - Mamba-2 + MoE + Attention hybrid with Multi-Token Prediction (MTP) Context Length Up to 1M tokens Minimum GPU Requirement 1× B200 OR 1× DGX Spark Supported Languages English, French, German, Italian, Japanese,…

Reasoning 120.0B ★ 13.0 ↓ 678.6K
NVIDIA-Nemotron-3-Super-120B-A12B-FP8 logo
NVIDIA-Nemotron-3-Super-120B-A12B-FP8
nvidia

:--- :--- Total Parameters 120B (12B active) Architecture LatentMoE - Mamba-2 + MoE + Attention hybrid with Multi-Token Prediction (MTP) Context Length Up to 1M tokens Minimum GPU Requirement 2× H100-80GB Supported Languages English, French, German, Italian, Japanese, Spanish, Ch…

Reasoning 120.0B ★ 13.0 ↓ 64K
command-a-vision-07-2025 logo
command-a-vision-07-2025
CohereLabs

Multimodal 112.0B ★ 13.0 ↓ 19K
granite-4.2-8b logo
granite-4.2-8b
ibm-granite

--- --- Developers Granite Team, IBM Model Type Decoder-only Dense Transformer (Reasoning) Architecture GraniteForCausalLM Base Model Granite-4.1-8B-Base Parameters 8B Context Length Natively Supports 128K (Long-context extension to 512K) Precision bfloat16 Tested Languages Engli…

Open Source 8.0B ★ 11.0 ↓ 133.6K
NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4 logo
NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4
nvidia

The post-training data has a cutoff date of November 28, 2025\. The pre-training data has a cutoff date of June 25, 2025\.

Open Source 30.0B ★ 9.0 ↓ 1.5M
NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 logo
NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
nvidia

The post-training data has a cutoff date of November 28, 2025\. The pre-training data has a cutoff date of June 25, 2025\.

Open Source 30.0B ★ 9.0 ↓ 905.2K
Llama-3.3-70B-Instruct logo
Llama-3.3-70B-Instruct
meta-llama

Open Source 70.0B ★ 9.0 ↓ 815.6K