AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,307 models for "Transformer" Compare
🤖
DialoGPT-medium
microsoft

A State-of-the-Art Large-scale Pretrained Response generation model (DialoGPT)

Open Source 345.0B ↓ 176.4K
🤖
GOT-OCR-2.0-hf
stepfun-ai

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model - HF Transformers 🤗 implementation

Multimodal 0.6B ↓ 167.3K
🤖
Kimi-K2-Instruct
moonshotai

📰   Tech Blog         📄   Paper

Open Source ↓ 166.4K
🤖
SmolVLM2-2.2B-Instruct
HuggingFaceTB

SmolVLM2-2.2B is a lightweight multimodal model designed to analyze video content. The model processes videos, images, and text inputs to generate text outputs - whether answering questions about media files, comparing visual content, or transcribing text from images. Despite its…

Multimodal 2.2B ↓ 162.7K
🤖
Kimi-K2.6-NVFP4
nvidia

Description: The NVIDIA Kimi-K2.6-NVFP4 model is the quantized version of the Moonshot AI's Kimi-K2.6 model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Kimi-K2.6 NVFP4 model is qu…

Open Source ↓ 162.3K
🤖
Nemotron-Labs-Diffusion-8B
nvidia

Nemotron-Labs-Diffusion is a tri-mode language model that supports both AR decoding and diffusion-based parallel decoding by simply switching the attention pattern of the same model during inference. The synergy between these two modes enables a third mode, called self-speculatio…

Open Source 8.0B ↓ 160.2K
🤖
Kimi-K2.5-NVFP4
nvidia

Description: The NVIDIA Kimi-K2.5-NVFP4 model is the quantized version of the Moonshot AI's Kimi-K2.5 model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Kimi-K2.5 NVFP4 model is qu…

Open Source ↓ 156.1K
🤖
Step-3.5-Flash
stepfun-ai

Step 3.5 Flash (visit website) is our most capable open-source foundation model, engineered to deliver frontier reasoning and agentic capabilities with exceptional efficiency. Built on a sparse Mixture of Experts (MoE) architecture, it selectively activates only 11B of its 196B p…

Open Source ↓ 155.3K
🤖
granite-3.0-8b-instruct
ibm-granite

Model Summary: Granite-3.0-8B-Instruct is a 8B parameter model finetuned from Granite-3.0-8B-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets. This model is developed using a diverse set of techniques…

Open Source 8.1B ↓ 152.6K
🤖
OLMoE-1B-7B-0924
allenai

OLMoE-1B-7B is a Mixture-of-Experts LLM with 1B active and 7B total parameters released in September 2024 (0924). It yields state-of-the-art performance among models with a similar cost (1B) and is competitive with much larger models like Llama2-13B. OLMoE is 100% open-source.

Open Source 7.0B ↓ 151.2K
🤖
SmolLM-135M
HuggingFaceTB

1. Model Summary 2. Limitations 3. Training 4. License 5. Citation

Open Source ↓ 145.9K
🤖
InternVL3-1B-hf
OpenGVLab

InternVL3-1B Transformers 🤗 Implementation

Multimodal 0.94B ↓ 143.3K