AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,308 models for "Transformer" Compare
🤖
GLM-5.1
zai-org

👋 Join our WeChat or Discord community. 📖 Check out the GLM-5.1 blog and GLM-5 Technical report . 📍 Use GLM-5.1 API services on Z.ai API Platform. 🔜 GLM-5.1 will be available on chat.z.ai in the coming days.

Open Source ↓ 87.6K
🤖
udop-large
microsoft

The UDOP model was proposed in Unifying Vision, Text, and Layout for Universal Document Processing by Zineng Tang, Ziyi Yang, Guoxin Wang, Yuwei Fang, Yang Liu, Chenguang Zhu, Michael Zeng, Cha Zhang, Mohit Bansal.

Multimodal 0.794B ↓ 86.5K
🤖
Llama-Guard-4-12B
meta-llama

Open Source 12.0B ↓ 85.9K
🤖
GLM-4.5-Air-FP8
zai-org

👋 Join our Discord community. 📖 Check out the GLM-4.5 technical blog , technical report , and Zhipu AI technical documentation . Code on GitHub 📍 Use GLM-4.5 API services on Z.ai API Platform (Global) or Zhipu AI Open Platform (Mainland China) . 👉 One click to GLM-4.5 .

Open Source ↓ 83.9K
🤖
Phi-3-vision-128k-instruct
microsoft

🎉 Phi-3.5 : [[mini-instruct]](https://huggingface.co/microsoft/Phi-3.5-mini-instruct); [[MoE-instruct]](https://huggingface.co/microsoft/Phi-3.5-MoE-instruct) ; [[vision-instruct]](https://huggingface.co/microsoft/Phi-3.5-vision-instruct)

Multimodal ↓ 83.9K
🤖
DeepSeek-V3.2-Exp
deepseek-ai

We are excited to announce the official release of DeepSeek-V3.2-Exp, an experimental version of our model. As an intermediate step toward our next-generation architecture, V3.2-Exp builds upon V3.1-Terminus by introducing DeepSeek Sparse Attention—a sparse attention mechanism de…

Open Source ↓ 82.6K
🤖
LFM2-1.2B
LiquidAI

LFM2 is a new generation of hybrid models developed by Liquid AI, specifically designed for edge AI and on-device deployment. It sets a new standard in terms of quality, speed, and memory efficiency.

Open Source 1.2B ↓ 81.8K
🤖
InternVL3_5-1B
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 1.1B ↓ 81.1K
🤖
Qwen3-Swallow-32B-RL-v0.2-AWQ-INT4
tokyotech-llm

Qwen3-Swallow v0.2 is a family of large language models available in 8B , 30B-A3B , and 32B parameter sizes. Built as bilingual Japanese-English models, they were developed through Continual Pre-Training (CPT), Supervised Fine-Tuning (SFT), and Reinforcement Learning with Verifia…

Reasoning 32.0B ↓ 78K
🤖
InternVL3-8B-hf
OpenGVLab

InternVL3-8B Transformers 🤗 Implementation

Multimodal 8.0B ↓ 76.9K
🤖
MiniCPM-V-4_5-AWQ
openbmb

A GPT-4o Level MLLM for Single Image, Multi Image and Video Understanding on Your Phone

Multimodal 8.7B ↓ 74.3K
🤖
internlm2-chat-7b
internlm

💻Github Repo • 🤔Reporting Issues • 📜Technical Report

Open Source 7.0B ↓ 74.1K