AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

440 models for "Base Model" Compare
Baichuan-M2-32B logo
Baichuan-M2-32B
baichuan-inc

This repository contains the model presented in Baichuan-M2: Scaling Medical Capability with Large Verifier System.

Open Source 32.0B ↓ 1.3K
Qwen-SEA-LION-v4-4B-VL logo
Qwen-SEA-LION-v4-4B-VL
aisingapore

SEA-LION is a collection of Large Language Models (LLMs) which have been pretrained and instruct-tuned for the Southeast Asia (SEA) region.

Multimodal 4.0B ↓ 1.1K
MiMo-7B-SFT logo
MiMo-7B-SFT
XiaomiMiMo

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ Unlocking the Reasoning Potential of Language Model From Pretraining to Posttraining ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

Open Source 7.0B ↓ 1.1K
Meta-Llama-3-70B logo
Meta-Llama-3-70B
NousResearch

Meta developed and released the Meta Llama 3 family of large language models (LLMs), a collection of pretrained and instruction tuned generative text models in 8 and 70B sizes. The Llama 3 instruction tuned models are optimized for dialogue use cases and outperform many of the av…

Open Source 70.6B ↓ 1.1K
RedPajama-INCITE-7B-Base logo
RedPajama-INCITE-7B-Base
togethercomputer

RedPajama-INCITE-7B-Base was developed by Together and leaders from the open-source AI community including Ontocord.ai, ETH DS3Lab, AAI CERC, Université de Montréal, MILA - Québec AI Institute, Stanford Center for Research on Foundation Models (CRFM), Stanford Hazy Research resea…

Open Source 7.0B ↓ 1.1K
xgen-7b-8k-base logo
xgen-7b-8k-base
Salesforce

Official research release for the family of XGen models ( 7B ) by Salesforce AI Research:

Open Source 7.0B ↓ 1K
Baichuan2-13B-Base logo
Baichuan2-13B-Base
baichuan-inc

🦉GitHub 💬WeChat 百川API支持搜索增强和192K长窗口,新增百川搜索增强知识库、限时免费! 🚀 百川大模型在线对话平台 已正式向公众开放 🎉

Open Source 13.0B ↓ 1K
Llama-3.1-Swallow-8B-Instruct-v0.2 logo
Llama-3.1-Swallow-8B-Instruct-v0.2
tokyotech-llm

Llama 3.1 Swallow is a series of large language models (8B, 70B) that were built by continual pre-training on the Meta Llama 3.1 models. Llama 3.1 Swallow enhanced the Japanese language capabilities of the original Llama 3.1 while retaining the English language capabilities. We u…

Open Source 8.0B ↓ 992
RedPajama-INCITE-Chat-3B-v1 logo
RedPajama-INCITE-Chat-3B-v1
togethercomputer

RedPajama-INCITE-Chat-3B-v1 was developed by Together and leaders from the open-source AI community including Ontocord.ai, ETH DS3Lab, AAI CERC, Université de Montréal, MILA - Québec AI Institute, Stanford Center for Research on Foundation Models (CRFM), Stanford Hazy Research re…

Open Source 3.0B ↓ 976
CapRL-Qwen3VL-4B logo
CapRL-Qwen3VL-4B
internlm

CapRL 📖 Paper 🏠 Github 🤗 CapRL Collection 🤗 Daily Paper

Multimodal 4.0B ↓ 931
nanowhale-100m-base logo
nanowhale-100m-base
HuggingFaceTB

A small ~110M parameter language model implementing the DeepSeek-V4 architecture from scratch. This is the pretrained base model — see HuggingFaceTB/nanowhale-100m for the SFT/chat version.

Open Source ↓ 930
stablelm-2-12b logo
stablelm-2-12b
stabilityai

Stable LM 2 12B is a 12.1 billion parameter decoder-only language model pre-trained on 2 trillion tokens of diverse multilingual and code datasets for two epochs.

Open Source 12.0B ↓ 922