AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,108 models for "Chat" Compare
🤖
gemma-4-26B-A4B-it-qat-q4_0-unquantized
google

Hugging Face GitHub Launch Blog Documentation Technical Report License : Apache 2.0 Authors : Google DeepMind

Open Source 26.0B ★ 26.0 ↓ 239.8K
🤖
NVIDIA-Nemotron-3-Super-120B-A12B-FP8
nvidia

:--- :--- Total Parameters 120B (12B active) Architecture LatentMoE - Mamba-2 + MoE + Attention hybrid with Multi-Token Prediction (MTP) Context Length Up to 1M tokens Minimum GPU Requirement 2× H100-80GB Supported Languages English, French, German, Italian, Japanese, Spanish, Ch…

Reasoning 120.0B ★ 26.0 ↓ 125.6K
🤖
MiMo-V2-Flash
XiaomiMiMo

🤗 HuggingFace   📔 Technical Report   📰 Blog   Play around!   🗨️ Xiaomi MiMo Studio   🎨 Xiaomi MiMo API Platform

Open Source ★ 25.0 ↓ 88K
🤖
Ling-3.0-tiny
inclusionAI

🤗 Hugging Face      🤖 ModelScope      🐙 OpenRouter   

Reasoning 7.9B ★ 25.0 ↓ 21.6K
🤖
MiMo-V2-Flash-Base
XiaomiMiMo

🤗 HuggingFace   📔 Technical Report   📰 Blog   Play around!   🗨️ Xiaomi MiMo Studio   🎨 Xiaomi MiMo API Platform

Open Source ★ 25.0 ↓ 354
🤖
gpt-oss-120b
openai

Try gpt-oss · Guides · Model card · OpenAI blog

Open Source 120.0B ★ 24.0 ↓ 5.4M
🤖
Qwen3.5-35B-A3B
Qwen

[!Note] This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.

Open Source 35.0B ★ 24.0 ↓ 2.5M
🤖
Qwen3.5-35B-A3B-FP8
Qwen

[!Note] This repository contains FP8-quantized model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc. The quantization method is fin…

Multimodal 35.0B ★ 24.0 ↓ 1.4M
🤖
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4
nvidia

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4

Open Source 30.0B ★ 24.0 ↓ 946.3K
🤖
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
nvidia

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Open Source 30.0B ★ 24.0 ↓ 438.3K
🤖
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4-DSpark
nvidia

The NVIDIA Nemotron-3.5-Lightning-30B-A3B-NVFP4-DSpark model is the DSpark speculative decoding checkpoint for NVIDIA's Nemotron-3.5-Lightning-30B-A3B model family, which is a hybrid LatentMoE language model designed for reasoning, chat, and agentic workflows. For more informatio…

Open Source 30.0B ★ 24.0 ↓ 200K
🤖
granite-4.2-30b
ibm-granite

--- --- Developers Granite Team, IBM Model Type Decoder-only Dense Transformer (Reasoning) Architecture GraniteForCausalLM Base Model Granite-4.1-30B-Base Parameters 30B Context Length Natively Supports 128K (Long-context extension to 512K) Precision bfloat16 Tested Languages Eng…

Open Source 30.0B ★ 24.0 ↓ 4.2K