AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

362 models for "Fine-tuned" Compare
NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 logo
NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
nvidia

:--- :--- Total Parameters 550B (55B active) Architecture LatentMoE - Mamba-2 + MoE + Attention hybrid with Multi-Token Prediction (MTP) Context Length Up to 1M tokens Minimum GPU Requirement 8x GB200/B200/GB300/B300, 16x H100, 8x H200 Supported Languages English, French, Spanish…

Open Source 550.0B ★ 23.0 ↓ 683.5K
NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4 logo
NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4
nvidia

:--- :--- Total Parameters 550B (55B active) Architecture LatentMoE - Mamba-2 + MoE + Attention hybrid with Multi-Token Prediction (MTP) Context Length Up to 1M tokens Minimum GPU Requirement 4xGB200, 4xB200, 4x GB300, 4x B300, 8xH100 Supported Languages English, French, Spanish,…

Open Source 550.0B ★ 23.0 ↓ 159.4K
Olmo-3-1025-7B logo
Olmo-3-1025-7B
allenai

We introduce Olmo 3, a new family of 7B and 32B models. This suite includes Base, Instruct, and Think variants. The Base models were trained using a staged training approach.

Open Source 7.0B ★ 20.0 ↓ 127.5K
Olmo-3-32B-Think-SFT logo
Olmo-3-32B-Think-SFT
allenai

We introduce Olmo 3, a new family of 7B and 32B models both Instruct and Think variants. Long chain-of-thought thinking improves reasoning tasks like math and coding.

Open Source 32.0B ★ 20.0 ↓ 43.5K
Olmo-3-1125-32B logo
Olmo-3-1125-32B
allenai

We introduce Olmo 3, a new family of 7B and 32B models. This suite includes Base, Instruct, and Think variants. The Base models were trained using a staged training approach.

Open Source 32.0B ★ 20.0 ↓ 21.9K
Olmo-3-32B-Think logo
Olmo-3-32B-Think
allenai

We introduce Olmo 3, a new family of 7B and 32B models both Instruct and Think variants. Long chain-of-thought thinking improves reasoning tasks like math and coding.

Open Source 32.0B ★ 20.0 ↓ 14.6K
qwen2.5-bakeneko-32b-instruct logo
qwen2.5-bakeneko-32b-instruct
rinna

Qwen2.5 Bakeneko 32B Instruct (rinna/qwen2.5-bakeneko-32b-instruct)

Open Source 32.0B ★ 20.0 ↓ 144
deepseek-r1-distill-qwen2.5-bakeneko-32b logo
deepseek-r1-distill-qwen2.5-bakeneko-32b
rinna

DeepSeek R1 Distill Qwen2.5 Bakeneko 32B (rinna/deepseek-r1-distill-qwen2.5-bakeneko-32b)

Reasoning 32.0B ★ 20.0 ↓ 137
qwq-bakeneko-32b logo
qwq-bakeneko-32b
rinna

QwQ Bakeneko 32B (rinna/qwq-bakeneko-32b)

Open Source 32.0B ★ 20.0 ↓ 128
Qwen3.6-35B-A3B-NVFP4 logo
Qwen3.6-35B-A3B-NVFP4
nvidia

Description: The NVIDIA Qwen3.6-35B-A3B-NVFP4 model is the quantized version of Alibaba's Qwen3.6-35B-A3B model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Qwen3.6-35B-A3B-NVFP4 m…

Open Source 35.0B ★ 18.0 ↓ 5.3M
Qwen3.5-122B-A10B-NVFP4 logo
Qwen3.5-122B-A10B-NVFP4
nvidia

Description: The NVIDIA Qwen3.5-122B-A10B-NVFP4 model is the quantized version of Alibaba's Qwen3.5-122B-A10B model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Qwen3.5-122B-A10B N…

Open Source 122.0B ★ 18.0 ↓ 1M
gemma-4-26B-A4B-it logo
gemma-4-26B-A4B-it
google

Hugging Face GitHub Launch Blog Documentation Technical Report License : Apache 2.0 Authors : Google DeepMind

Open Source 26.0B ★ 17.0 ↓ 12.3M