AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,301 models for "Transformer" Compare
🤖
OpenELM-450M
apple

Sachin Mehta, Mohammad Hossein Sekhavat, Qingqing Cao, Maxwell Horton, Yanzi Jin, Chenfan Sun, Iman Mirzadeh, Mahyar Najibi, Dmitry Belenko, Peter Zatloukal, Mohammad Rastegari

Open Source ↓ 379
🤖
MiMo-7B-RL-Zero
XiaomiMiMo

━━━━━━━━━━━━━━━━━━━━━━━━━ Unlocking the Reasoning Potential of Language Model From Pretraining to Posttraining ━━━━━━━━━━━━━━━━━━━━━━━━━

Open Source 7.0B ↓ 378
🤖
japanese-stablelm-instruct-gamma-7b
stabilityai

This is a 7B-parameter decoder-only Japanese language model fine-tuned on instruction-following datasets, built on top of the base model Japanese Stable LM Base Gamma 7B.

Open Source 7.0B ↓ 377
🤖
Yi-6B-Chat-8bits
01-ai

Building the Next Generation of Open-Source and Bilingual LLMs

Open Source 6.0B ↓ 373
🤖
MiniCPM-V-4_5-GPTQ
openbmb

A GPT-4o Level MLLM for Single Image, Multi Image and High-FPS Video Understanding on Your Phone

Open Source ↓ 372
🤖
Llama-SEA-LION-v3-8B
aisingapore

SEA-LION is a collection of Large Language Models (LLMs) which have been pretrained and instruct-tuned for the Southeast Asia (SEA) region.

Open Source 8.0B ↓ 371
🤖
HiLS-Attention-7B
tencent

HiLS-Attention is a chunk-wise sparse attention mechanism that learns chunk selection end-to-end under the language-modeling loss, enabling native sparse training for efficient long-context modeling. This repository hosts the 7B checkpoint continued-trained on top of an OLMo3-sty…

Open Source 7.0B ↓ 365
🤖
cogagent-chat-hf
zai-org

🔥 News : The new version CogAgent-9B-20241220 has been released! Welcome to visit CogAgent GitHub and Technical Report to explore and use our latest model.

Open Source ↓ 364
🤖
Llama-SEA-LION-v2-8B-IT
aisingapore

Llama-SEA-LION-v2-8B-IT SEA-LION is a collection of Large Language Models (LLMs) which have been pretrained and instruct-tuned for the Southeast Asia (SEA) region.

Open Source 8.0B ↓ 363
🤖
ZwZ-4B
inclusionAI

ZwZ-4B is a fine-grained multimodal perception model built upon Qwen3-VL-4B. It is trained using Region-to-Image Distillation (R2I) combined with reinforcement learning, enabling superior fine-grained visual understanding in a single forward pass — no inference-time zooming or to…

Open Source 4.0B ↓ 362
🤖
stablelm-base-alpha-7b-v2
stabilityai

StableLM-Base-Alpha-7B-v2 is a 7 billion parameter decoder-only language model pre-trained on diverse English datasets. This model is the successor to the first StableLM-Base-Alpha-7B model, addressing previous shortcomings through the use of improved data sources and mixture rat…

Open Source 7.0B ↓ 357
🤖
SmolLM3-3B-GSM8K-SFT
HuggingFaceTB

Fine-tuned version of HuggingFaceTB/SmolLM3-3B-Base optimized for grade school math (GSM8K benchmark).

Open Source 3.0B ↓ 355