AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

362 models for "Fine-tuned" Compare
gemma-4-26B-A4B-it-qat-q4_0-unquantized logo
gemma-4-26B-A4B-it-qat-q4_0-unquantized
google

Hugging Face GitHub Launch Blog Documentation Technical Report License : Apache 2.0 Authors : Google DeepMind

Open Source 26.0B ★ 17.0 ↓ 125.1K
gemma-4-31B-it logo
gemma-4-31B-it
google

Hugging Face GitHub Launch Blog Documentation Technical Report License : Apache 2.0 Authors : Google DeepMind

Open Source 31.0B ★ 15.0 ↓ 9.6M
gemma-4-31B logo
gemma-4-31B
google

Hugging Face GitHub Launch Blog Documentation Technical Report License : Apache 2.0 Authors : Google DeepMind

Open Source 31.0B ★ 15.0 ↓ 555K
gemma-4-31B-it-qat-q4_0-unquantized logo
gemma-4-31B-it-qat-q4_0-unquantized
google

Hugging Face GitHub Launch Blog Documentation Technical Report License : Apache 2.0 Authors : Google DeepMind

Open Source 31.0B ★ 15.0 ↓ 315.7K
gemma-4-31B-it-qat-w4a16-ct logo
gemma-4-31B-it-qat-w4a16-ct
google

Hugging Face GitHub Launch Blog Documentation Technical Report License : Apache 2.0 Authors : Google DeepMind

Open Source 31.0B ★ 15.0 ↓ 249.5K
NVIDIA-Nemotron-3-Super-120B-A12B-BF16 logo
NVIDIA-Nemotron-3-Super-120B-A12B-BF16
nvidia

:--- :--- Total Parameters 120B (12B active) Architecture LatentMoE - Mamba-2 + MoE + Attention hybrid with Multi-Token Prediction (MTP) Context Length Up to 1M tokens Minimum GPU Requirement 8× H100-80GB Supported Languages English, French, German, Italian, Japanese, Spanish, Ch…

Reasoning 120.0B ★ 13.0 ↓ 1M
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 logo
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4
nvidia

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4

Open Source 30.0B ★ 13.0 ↓ 839.9K
NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 logo
NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4
nvidia

:--- :--- Total Parameters 120B (12B active) Architecture LatentMoE - Mamba-2 + MoE + Attention hybrid with Multi-Token Prediction (MTP) Context Length Up to 1M tokens Minimum GPU Requirement 1× B200 OR 1× DGX Spark Supported Languages English, French, German, Italian, Japanese,…

Reasoning 120.0B ★ 13.0 ↓ 678.6K
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 logo
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
nvidia

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Open Source 30.0B ★ 13.0 ↓ 599K
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4-DSpark logo
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4-DSpark
nvidia

The NVIDIA Nemotron-3.5-Lightning-30B-A3B-NVFP4-DSpark model is the DSpark speculative decoding checkpoint for NVIDIA's Nemotron-3.5-Lightning-30B-A3B model family, which is a hybrid LatentMoE language model designed for reasoning, chat, and agentic workflows. For more informatio…

Open Source 30.0B ★ 13.0 ↓ 134.1K
NVIDIA-Nemotron-3-Super-120B-A12B-FP8 logo
NVIDIA-Nemotron-3-Super-120B-A12B-FP8
nvidia

:--- :--- Total Parameters 120B (12B active) Architecture LatentMoE - Mamba-2 + MoE + Attention hybrid with Multi-Token Prediction (MTP) Context Length Up to 1M tokens Minimum GPU Requirement 2× H100-80GB Supported Languages English, French, German, Italian, Japanese, Spanish, Ch…

Reasoning 120.0B ★ 13.0 ↓ 64K
gpt-oss-120b logo
gpt-oss-120b
openai

Try gpt-oss · Guides · Model card · OpenAI blog

Open Source 120.0B ★ 12.0 ↓ 4.1M