AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

440 models for "Base Model" Compare
LFM2.5-VL-1.6B logo
LFM2.5-VL-1.6B
LiquidAI

LFM2.5‑VL-1.6B is Liquid AI's refreshed version of the first vision-language model, LFM2-VL-1.6B, built on an updated backbone LFM2.5-1.2B-Base and tuned for stronger real-world performance. Find more about LFM2.5 family of models in our blog post.

Multimodal 1.6B ★ 5.0 ↓ 30.9K
Olmo-3-7B-Instruct-DPO logo
Olmo-3-7B-Instruct-DPO
allenai

We introduce Olmo 3, a new family of 7B and 32B models both Instruct and Think variants. Long chain-of-thought thinking improves reasoning tasks like math and coding.

Open Source 7.0B ★ 5.0 ↓ 14.7K
LFM2.5-1.2B-Thinking logo
LFM2.5-1.2B-Thinking
LiquidAI

LFM2.5 is a new family of hybrid models designed for on-device deployment . It builds on the LFM2 architecture with extended pre-training and reinforcement learning.

Reasoning 1.2B ★ 5.0 ↓ 8.5K
granite-4.0-micro-base logo
granite-4.0-micro-base
ibm-granite

Model Summary: Granite-4.0-Micro-Base is a decoder-only, long-context language model designed for a wide range of text-to-text generation tasks. It also supports Fill-in-the-Middle (FIM) code completion through the use of specialized prefix and suffix tokens. The model is trained…

Open Source 3.0B ★ 5.0 ↓ 4.3K
granite-4.1-3b-base logo
granite-4.1-3b-base
ibm-granite

Model Summary: Granite‑4.1‑3B‑Base is a decoder‑only language model with long‑context capabilities, designed to support a broad range of general text‑to‑text generation tasks, as well as fill‑in‑the‑Middle (FIM) code completion. This model shares the same underlying architecture…

Open Source 3.0B ★ 4.0 ↓ 14.6K
Llama-3.1-8B logo
Llama-3.1-8B
meta-llama

Open Source 8.0B ★ 2.0 ↓ 513.5K
granite-4.0-1b-base logo
granite-4.0-1b-base
ibm-granite

Model Summary: Granite-4.0-1B-Base is a lightweight decoder-only language model designed for scenarios where efficiency and speed are critical. They can run on resource-constrained devices such as smartphones or IoT hardware, enabling offline and privacy-preserving applications.…

Open Source 1.6B ★ 2.0 ↓ 44.7K
LFM2-2.6B-Longevity logo
LFM2-2.6B-Longevity
LiquidAI

Longevity-LLM - LFM2-2.6B Longevity-LLM (L-LLM) is a family of compact, domain-adapted language models for interpreting heterogeneous aging biology data. Longevity LFMs are available in two sizes:

Open Source 2.6B ★ 2.0 ↓ 2.1K
Mistral-7B-Instruct-v0.2 logo
Mistral-7B-Instruct-v0.2
mistralai

py from mistral common.tokens.tokenizers.mistral import MistralTokenizer from mistral common.protocol.instruct.messages import UserMessage from mistral common.protocol.instruct.request import ChatCompletionRequest

Open Source 7.0B ↓ 1.6M
DeepSeek-V3 logo
DeepSeek-V3
deepseek-ai

We present DeepSeek-V3, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token. To achieve efficient inference and cost-effective training, DeepSeek-V3 adopts Multi-head Latent Attention (MLA) and DeepSeekMoE architectures, w…

Open Source ↓ 1.4M
DeepSeek-R1 logo
DeepSeek-R1
deepseek-ai

We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With R…

Reasoning ↓ 1.1M
DeepSeek-R1-Distill-Qwen-1.5B logo
DeepSeek-R1-Distill-Qwen-1.5B
deepseek-ai

We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With R…

Reasoning 1.5B ↓ 1M