AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,308 models for "Transformer" Compare
🤖
GLM-4.6V-Flash
zai-org

This model is part of the GLM-V family of models, introduced in the paper GLM-4.1V-Thinking and GLM-4.5V: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

Open Source ↓ 98.5K
🤖
GLM-4.5
zai-org

👋 Join our Discord community. 📖 Check out the GLM-4.5 technical blog , technical report , and Zhipu AI technical documentation . 📍 Use GLM-4.5 API services on Z.ai API Platform (Global) or Zhipu AI Open Platform (Mainland China) . 👉 One click to GLM-4.5 .

Open Source ↓ 95.8K
🤖
Nemotron-Labs-Diffusion-8B-Base
nvidia

Nemotron-Labs-Diffusion is a tri-mode language model that supports both AR decoding and diffusion-based parallel decoding by simply switching the attention pattern of the same model during inference. The synergy between these two modes enables a third mode, called self-speculatio…

Open Source 8.0B ↓ 95.5K
🤖
LFM2.5-350M
LiquidAI

LFM2.5 is a new family of hybrid models designed for on-device deployment . It builds on the LFM2 architecture with extended pre-training and reinforcement learning.

Open Source ↓ 94.2K
🤖
DeepSeek-R1-Distill-Llama-70B
deepseek-ai

We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With R…

Reasoning 70.0B ↓ 94K
🤖
Qwen3-14B-NVFP4
nvidia

Description: The NVIDIA Qwen3-14B FP4 model is the quantized version of Alibaba's Qwen3-14B model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Qwen3-14B FP4 model is quantized with…

Open Source 14.8B ↓ 92.1K
🤖
xglm-564M
facebook

XGLM-564M is a multilingual autoregressive language model (with 564 million parameters) trained on a balanced corpus of a diverse set of 30 languages totaling 500 billion sub-tokens. It was introduced in the paper Few-shot Learning with Multilingual Language Models by Xi Victoria…

Open Source ↓ 92K
🤖
internlm3-8b-instruct
internlm

💻Github Repo • 🤗Demo • 🤔Reporting Issues • 📜Technical Report

Open Source 8.0B ↓ 89.2K
🤖
sarashina2.2-0.5b-instruct-v0.1
sbintuitions

sbintuitions/sarashina2.2-0.5b-instruct-v0.1

Open Source 0.5B ↓ 88.8K
🤖
granite-4.0-h-tiny
ibm-granite

📣 Update [10-07-2025]: Added a default system prompt to the chat template to guide the model towards more professional, accurate, and safe responses.

Open Source 7.0B ↓ 88.6K
🤖
Llama-3.1-70B-Instruct-FP8
nvidia

Description: The NVIDIA Llama 3.1 70B Instruct FP8 model is the quantized version of the Meta's Llama 3.1 70B Instruct model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Llama 3.1…

Open Source 70.0B ↓ 88.6K
🤖
DeepSeek-R1-0528-NVFP4-v2
nvidia

Description: The NVIDIA DeepSeek-R1-0528-FP4 v2 model is the quantized version of the DeepSeek AI's DeepSeek R1 0528 model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA DeepSeek R1…

Reasoning ↓ 87.6K