AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

440 models for "Base Model" Compare
LFM2-VL-1.6B logo
LFM2-VL-1.6B
LiquidAI

LFM2‑VL is Liquid AI's first series of multimodal models, designed to process text and images with variable resolutions. Built on the LFM2 backbone, it is optimized for low-latency and edge AI applications.

Multimodal 1.6B ↓ 66.2K
Meta-Llama-3-8B logo
Meta-Llama-3-8B
NousResearch

Meta developed and released the Meta Llama 3 family of large language models (LLMs), a collection of pretrained and instruction tuned generative text models in 8 and 70B sizes. The Llama 3 instruction tuned models are optimized for dialogue use cases and outperform many of the av…

Open Source 8.0B ↓ 64.5K
Phi-mini-MoE-instruct logo
Phi-mini-MoE-instruct
microsoft

Phi-mini-MoE is a lightweight Mixture of Experts (MoE) model with 7.6B total parameters and 2.4B activated parameters. It is compressed and distilled from the base model shared by Phi-3.5-MoE and GRIN-MoE using the SlimMoE approach, then post-trained via supervised fine-tuning an…

Open Source ↓ 64K
LFM2.5-350M logo
LFM2.5-350M
LiquidAI

LFM2.5 is a new family of hybrid models designed for on-device deployment . It builds on the LFM2 architecture with extended pre-training and reinforcement learning.

Open Source ↓ 63.5K
DeepSeek-R1-Distill-Llama-70B logo
DeepSeek-R1-Distill-Llama-70B
deepseek-ai

We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With R…

Reasoning 70.0B ↓ 61.9K
internlm-chat-7b logo
internlm-chat-7b
internlm

InternLM has open-sourced a 7 billion parameter base model and a chat model tailored for practical scenarios. The model has the following characteristics: - It leverages trillions of high-quality tokens for training to establish a powerful knowledge base. - It supports an 8k cont…

Open Source 7.0B ↓ 61.5K
Phi-tiny-MoE-instruct logo
Phi-tiny-MoE-instruct
microsoft

Phi-tiny-MoE is a lightweight Mixture of Experts (MoE) model with 3.8B total parameters and 1.1B activated parameters. It is compressed and distilled from the base model shared by Phi-3.5-MoE and GRIN-MoE using the SlimMoE approach, then post-trained via supervised fine-tuning an…

Open Source ↓ 61K
phi-1_5 logo
phi-1_5
microsoft

The language model Phi-1.5 is a Transformer with 1.3 billion parameters. It was trained using the same data sources as phi-1, augmented with a new data source that consists of various NLP synthetic texts. When assessed against benchmarks testing common sense, language understandi…

Open Source 1.3B ↓ 60.5K
granite-4.0-3b-vision logo
granite-4.0-3b-vision
ibm-granite

Model Summary: Granite-4.0-3B-Vision is a vision-language model (VLM) designed for enterprise-grade document data extraction. It focuses on specialized, complex extraction tasks that ultracompact models often struggle with:

Multimodal 4.0B ↓ 56.3K
Phi-4-mini-reasoning logo
Phi-4-mini-reasoning
microsoft

Phi-4-mini-reasoning is a lightweight open model built upon synthetic data with a focus on high-quality, reasoning dense data further finetuned for more advanced math reasoning capabilities. The model belongs to the Phi-4 model family and supports 128K token context length.

Open Source ↓ 56.2K
GLM-4.5-Air-FP8 logo
GLM-4.5-Air-FP8
zai-org

👋 Join our Discord community. 📖 Check out the GLM-4.5 technical blog , technical report , and Zhipu AI technical documentation . Code on GitHub 📍 Use GLM-4.5 API services on Z.ai API Platform (Global) or Zhipu AI Open Platform (Mainland China) . 👉 One click to GLM-4.5 .

Open Source ↓ 55.4K
Meta-Llama-3-70B-Instruct logo
Meta-Llama-3-70B-Instruct
NousResearch

Meta developed and released the Meta Llama 3 family of large language models (LLMs), a collection of pretrained and instruction tuned generative text models in 8 and 70B sizes. The Llama 3 instruction tuned models are optimized for dialogue use cases and outperform many of the av…

Open Source 70.0B ↓ 52.6K