AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

440 models for "Base Model" Compare
HunyuanOCR logo
HunyuanOCR
tencent

HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better

Multimodal 1.0B ↓ 939.5K
Qwen2-0.5B logo
Qwen2-0.5B
Qwen

Qwen2 is the new series of Qwen large language models. For Qwen2, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters, including a Mixture-of-Experts model. This repo contains the 0.5B Qwen2 base language mod…

Open Source 0.5B ↓ 830.7K
Qwen2.5-Math-1.5B logo
Qwen2.5-Math-1.5B
Qwen

[!Warning] 🚨 Qwen2.5-Math mainly supports solving English and Chinese math problems through CoT and TIR. We do not recommend using this series of models for other tasks.

Reasoning 1.5B ↓ 788.4K
phi-2 logo
phi-2
microsoft

Phi-2 is a Transformer with 2.7 billion parameters. It was trained using the same data sources as Phi-1.5, augmented with a new data source that consists of various NLP synthetic texts and filtered websites (for safety and educational value). When assessed against benchmarks test…

Open Source ↓ 621.8K
GLM-4.5-Air logo
GLM-4.5-Air
zai-org

👋 Join our Discord community. 📖 Check out the GLM-4.5 technical blog , technical report , and Zhipu AI technical documentation . 📍 Use GLM-4.5 API services on Z.ai API Platform (Global) or Zhipu AI Open Platform (Mainland China) . 👉 One click to GLM-4.5 .

Open Source ↓ 502.5K
bloom-560m logo
bloom-560m
bigscience

BLOOM LM BigScience Large Open-science Open-access Multilingual Language Model Model Card

Open Source 0.56B ↓ 459.2K
Mistral-7B-v0.1 logo
Mistral-7B-v0.1
mistralai

The Mistral-7B-v0.1 Large Language Model (LLM) is a pretrained generative text model with 7 billion parameters. Mistral-7B-v0.1 outperforms Llama 2 13B on all benchmarks we tested.

Open Source 7.0B ↓ 433K
DeepSeek-R1-Distill-Qwen-32B logo
DeepSeek-R1-Distill-Qwen-32B
deepseek-ai

We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With R…

Reasoning 32.0B ↓ 423.3K
Llama-2-7b-chat-hf logo
Llama-2-7b-chat-hf
meta-llama

Open Source 7.0B ↓ 347.5K
Kimi-K2-Instruct logo
Kimi-K2-Instruct
moonshotai

📰   Tech Blog         📄   Paper

Open Source ↓ 338.1K
DeepSeek-R1-Distill-Qwen-14B logo
DeepSeek-R1-Distill-Qwen-14B
deepseek-ai

We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With R…

Reasoning 14.0B ↓ 332K
OLMo-2-0425-1B logo
OLMo-2-0425-1B
allenai

We introduce OLMo 2 1B, the smallest model in the OLMo 2 family. OLMo 2 was pre-trained on OLMo-mix-1124 and uses Dolmino-mix-1124 for mid-training.

Open Source 1.0B ↓ 314.9K