AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,302 models for "Transformer" Compare
🤖
Mistral-7B-Instruct-v0.2
mistralai

py from mistral common.tokens.tokenizers.mistral import MistralTokenizer from mistral common.protocol.instruct.messages import UserMessage from mistral common.protocol.instruct.request import ChatCompletionRequest

Open Source 7.0B ↓ 1.2M
🤖
MiniMax-M2.7
MiniMaxAI

Join Our 💬 WeChat 🧩 Discord community. MiniMax Agent ⚡️ API CLI MiniMax Website 🤗 Hugging Face 🐙 GitHub 🤖️ ModelScope 📄 LICENSE

Open Source ↓ 1.2M
🤖
Qwen3-4B-Instruct-2507-FP8
Qwen

We introduce the updated version of the Qwen3-4B-FP8 non-thinking mode , named Qwen3-4B-Instruct-2507-FP8 , featuring the following key enhancements:

Open Source 4.0B ↓ 1.1M
🤖
Qwen3-Coder-30B-A3B-Instruct-FP8
Qwen

Qwen3-Coder is available in multiple sizes. Today, we're excited to introduce Qwen3-Coder-30B-A3B-Instruct-FP8 . This streamlined model maintains impressive performance and efficiency, featuring the following key enhancements:

Code 30.5B ↓ 1.1M
🤖
DeepSeek-V3-0324
deepseek-ai

DeepSeek-V3-0324 demonstrates notable improvements over its predecessor, DeepSeek-V3, in several key aspects.

Open Source ↓ 1.1M
🤖
DeepSeek-V3
deepseek-ai

We present DeepSeek-V3, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token. To achieve efficient inference and cost-effective training, DeepSeek-V3 adopts Multi-head Latent Attention (MLA) and DeepSeekMoE architectures, w…

Open Source ↓ 1.1M
🤖
DeepSeek-OCR-2
deepseek-ai

🌟 Github 📥 Model Download 📄 Paper Link 📄 Arxiv Paper Link DeepSeek-OCR 2: Visual Causal Flow Explore more human-like visual encoding.

Multimodal 3.0B ↓ 1.1M
🤖
Qwen2.5-1.5B
Qwen

Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2:

Open Source 1.5B ↓ 1M
🤖
OLMo-2-0425-1B
allenai

We introduce OLMo 2 1B, the smallest model in the OLMo 2 family. OLMo 2 was pre-trained on OLMo-mix-1124 and uses Dolmino-mix-1124 for mid-training.

Open Source 1.0B ↓ 1M
🤖
medgemma-4b-it
google

Multimodal 4.0B ↓ 1M
🤖
Qwen3-0.6B-Base
Qwen

Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Building upon extensive advancements in training data, model architecture, and optimization techniques, Qwen3 delivers the followin…

Open Source 0.6B ↓ 989.1K
🤖
Cosmos-Reason2-2B
nvidia

Multimodal 2.0B ↓ 982.2K