AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,294 models for "Transformer" Compare
🤖
Ling-3.0-flash-base-30T
inclusionAI

🤗 Hugging Face      🤖 ModelScope      🐙 OpenRouter   

Open Source ★ 38.0 ↓ 1.9K
🤖
Ling-3.0-flash-base
inclusionAI

🤗 Hugging Face      🤖 ModelScope      🐙 OpenRouter   

Open Source ★ 38.0 ↓ 1.9K
🤖
Ling-3.0-flash-base-midtrain
inclusionAI

🤗 Hugging Face      🤖 ModelScope      🐙 OpenRouter   

Open Source ★ 38.0 ↓ 1.8K
🤖
MiMo-V2.5-Base
XiaomiMiMo

🤗 HuggingFace   📰 Blog   🎨 Xiaomi MiMo API Platform   🗨️ Xiaomi MiMo Studio  

Open Source ★ 38.0 ↓ 344
🤖
NVIDIA: Nemotron 3 Ultra (free)
nvidia

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

Closed Source ★ 38.0
🤖
NVIDIA: Nemotron 3 Ultra (batch)
nvidia

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

Closed Source ★ 38.0
🤖
NVIDIA: Nemotron 3 Ultra
nvidia

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

Closed Source ★ 38.0
🤖
Solar-Open2-250B
upstage

Solar Open 2 is Upstage’s 250B-A15B open-weight large language model, built for agentic use cases such as office productivity, document-intensive work, and coding. Its Hybrid-Attention Mixture-of-Experts (MoE) architecture with linear attention delivers highly efficient inference…

Open Source 250.0B ★ 37.0 ↓ 24.9K
🤖
Qwen3.5-397B-A17B-NVFP4
nvidia

Description: The NVIDIA Qwen3.5-397B-A17B NVFP4 model is the quantized version of Alibaba's Qwen3.5-397B-A17B model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Qwen3.5-397B-A17B N…

Open Source 397.0B ★ 34.0 ↓ 186K
🤖
LongCat-2.0
meituan-longcat

Model Introduction We introduce LongCat-2.0, a large-scale MoE language model with 1.6 trillion total parameters and ~48 billion activated per token — a substantial step up from previous LongCat models, accompanied by several architectural improvements.

Open Source ★ 34.0 ↓ 1.5K
🤖
LongCat-2.0-FP8
meituan-longcat

Model Introduction We introduce LongCat-2.0, a large-scale MoE language model with 1.6 trillion total parameters and ~48 billion activated per token — a substantial step up from previous LongCat models, accompanied by several architectural improvements.

Open Source ★ 34.0 ↓ 202
🤖
LongCat-2.0-INT8
meituan-longcat

Model Introduction We introduce LongCat-2.0, a large-scale MoE language model with 1.6 trillion total parameters and ~48 billion activated per token — a substantial step up from previous LongCat models, accompanied by several architectural improvements.

Open Source ★ 34.0 ↓ 153