AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,295 models for "Transformer" Compare
🤖
NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
nvidia

The post-training data has a cutoff date of November 28, 2025\. The pre-training data has a cutoff date of June 25, 2025\.

Open Source 30.0B ★ 15.0 ↓ 871K
🤖
NVIDIA-Nemotron-3-Nano-30B-A3B-FP8
nvidia

The post-training data has a cutoff date of November 28, 2025\. The pre-training data has a cutoff date of June 25, 2025\.

Open Source 30.0B ★ 15.0 ↓ 685.6K
🤖
NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4
nvidia

The post-training data has a cutoff date of November 28, 2025\. The pre-training data has a cutoff date of June 25, 2025\.

Open Source 30.0B ★ 15.0 ↓ 580.3K
🤖
NVIDIA-Nemotron-3-Nano-30B-A3B-Base-BF16
nvidia

NVIDIA-Nemotron-3-Nano-30B-A3B-Base-BF16

Open Source 30.0B ★ 15.0 ↓ 121.6K
🤖
InternVL3_5-GPT-OSS-20B-A4B-Preview
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 20.0B ★ 15.0 ↓ 30.8K
🤖
Solar-Open-100B
upstage

Solar Open is Upstage's flagship 102B-parameter large language model, trained entirely from scratch and released under the Upstage Solar License (see LICENSE for details). As a Mixture-of-Experts (MoE) architecture, it delivers enterprise-grade performance in reasoning, instructi…

Open Source 100.0B ★ 15.0 ↓ 14.8K
🤖
granite-4.2-3b
ibm-granite

--- --- Developers Granite Team, IBM Model Type Decoder-only Dense Transformer (Reasoning) Architecture GraniteForCausalLM Base Model Granite-4.1-3B-Base Parameters 3B Context Length Natively Supports 128K (Long-context extension to 512K) Precision bfloat16 Tested Languages Engli…

Open Source 3.0B ★ 14.0 ↓ 8.2K
🤖
diffusiongemma-26B-A4B-it
google

Hugging Face GitHub Launch Blog Documentation License : Apache 2.0 Authors : Google DeepMind

Open Source 26.0B ★ 13.0 ↓ 1.3M
🤖
diffusiongemma-26B-A4B-it-NVFP4
nvidia

--- pipeline tag: text-generation base model: google/diffusiongemma-26B-A4B-it license: apache-2.0 license name: apache-license-2.0 license link: https://ai.google.dev/gemma/apache 2 tags: - nvidia - ModelOpt - DiffusionGemma-26B-A4B-IT - quantized - NVFP4 - nvfp4 ---

Open Source 26.0B ★ 13.0 ↓ 299.1K
🤖
MiniCPM5-1B
openbmb

MiniCPM Tech Report MiniCPM Wiki(Chinese) GitHub Repo UltraData MiniCPM Desk Pet Online Demo

Open Source 1.0B ★ 12.0 ↓ 790.5K
🤖
Llama-3_3-Nemotron-Super-49B-v1_5-FP8
nvidia

Llama-3.3-Nemotron-Super-49B-v1.5-FP8 is a significantly upgraded version of Llama-3.3-Nemotron-Super-49B-v1 and is a large language model (LLM) which is a derivative of Meta Llama-3.3-70B-Instruct (AKA the reference model). It is a reasoning model that is post trained for reason…

Reasoning 49.0B ★ 12.0 ↓ 253.4K
🤖
MiniCPM5-1B-SFT
openbmb

MiniCPM Tech Report MiniCPM Wiki(Chinese) GitHub Repo UltraData MiniCPM Desk Pet Online Demo

Open Source 1.0B ★ 12.0 ↓ 15.3K