AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

440 models for "Base Model" Compare
DeepSeek-R1-Distill-Qwen-7B logo
DeepSeek-R1-Distill-Qwen-7B
deepseek-ai

We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With R…

Reasoning 7.0B ↓ 292.8K
Kimi-K2.5-NVFP4 logo
Kimi-K2.5-NVFP4
nvidia

Description: The NVIDIA Kimi-K2.5-NVFP4 model is the quantized version of the Moonshot AI's Kimi-K2.5 model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Kimi-K2.5 NVFP4 model is qu…

Open Source ↓ 268.4K
DeepSeek-V2-Lite logo
DeepSeek-V2-Lite
deepseek-ai

Model Download Evaluation Results Model Architecture API Platform License Citation

Open Source ↓ 190K
DeepSeek-R1-Distill-Llama-8B logo
DeepSeek-R1-Distill-Llama-8B
deepseek-ai

We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With R…

Reasoning 8.0B ↓ 187.7K
GLM-4.1V-9B-Thinking logo
GLM-4.1V-9B-Thinking
zai-org

📖 View the GLM-4.1V-9B-Thinking paper . 📍 Using GLM-4.1V-9B-Thinking API at Zhipu Foundation Model Open Platform

Open Source 9.0B ↓ 164.2K
Kimi-K2.6-NVFP4 logo
Kimi-K2.6-NVFP4
nvidia

Description: The NVIDIA Kimi-K2.6-NVFP4 model is the quantized version of the Moonshot AI's Kimi-K2.6 model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Kimi-K2.6 NVFP4 model is qu…

Open Source ↓ 153.2K
Mistral-7B-Instruct-v0.1 logo
Mistral-7B-Instruct-v0.1
mistralai

py from mistral common.tokens.tokenizers.mistral import MistralTokenizer from mistral common.protocol.instruct.messages import UserMessage from mistral common.protocol.instruct.request import ChatCompletionRequest

Open Source 7.0B ↓ 152.3K
OLMoE-1B-7B-0125-Instruct logo
OLMoE-1B-7B-0125-Instruct
allenai

OLMoE-1B-7B-0125-Instruct January 2025 is post-trained variant of the OLMoE-1B-7B January 2025 model, which has undergone supervised finetuning on an OLMo-specific variant of the Tülu 3 dataset and further DPO training on this dataset, and finally RLVR training using this data. T…

Open Source 1.0B ↓ 134.8K
DeepSeek-V2-Lite-Chat logo
DeepSeek-V2-Lite-Chat
deepseek-ai

Model Download Evaluation Results Model Architecture API Platform License Citation

Open Source ↓ 134.4K
Meta-Llama-3-70B logo
Meta-Llama-3-70B
meta-llama

Open Source 70.0B ↓ 130K
gemma-3-4b-pt logo
gemma-3-4b-pt
google

Open Source 4.0B ↓ 129.3K
GLM-4.5 logo
GLM-4.5
zai-org

👋 Join our Discord community. 📖 Check out the GLM-4.5 technical blog , technical report , and Zhipu AI technical documentation . 📍 Use GLM-4.5 API services on Z.ai API Platform (Global) or Zhipu AI Open Platform (Mainland China) . 👉 One click to GLM-4.5 .

Open Source ↓ 124.4K