AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,111 models for "Chat" Compare
🤖
AquilaChat2-7B
BAAI

We opensource our Aquila2 series, now including Aquila2 , the base language models, namely Aquila2-7B and Aquila2-34B , as well as AquilaChat2 , the chat models, namely AquilaChat2-7B and AquilaChat2-34B , as well as the long-text chat models, namely AquilaChat2-7B-16k and Aquila…

Open Source 7.0B ↓ 712
🤖
DeepSeek-Prover-V2-671B
deepseek-ai

We introduce DeepSeek-Prover-V2, an open-source large language model designed for formal theorem proving in Lean 4, with initialization data collected through a recursive theorem proving pipeline powered by DeepSeek-V3. The cold-start training procedure begins by prompting DeepSe…

Open Source 671.0B ↓ 707
🤖
nanowhale-100m-base
HuggingFaceTB

A small ~110M parameter language model implementing the DeepSeek-V4 architecture from scratch. This is the pretrained base model — see HuggingFaceTB/nanowhale-100m for the SFT/chat version.

Open Source ↓ 706
🤖
DeepSeek-V2.5-1210
deepseek-ai

DeepSeek-V2.5-1210 is an upgraded version of DeepSeek-V2.5, with improvements across various capabilities:

Open Source ↓ 701
🤖
Ling-2.6-1T
inclusionAI

🤗 Hugging Face      🤖 ModelScope      🐙 OpenRouter       Tech Report

Open Source ↓ 699
🤖
SingGuard-NSFA-4B
inclusionAI

SingGuard-NSFA: Extensible Guardrails for Agentic AI via Generative Reasoning and Real-Time Classification

Open Source 4.0B ↓ 694
🤖
Llama-2-7B-32K-Instruct
togethercomputer

Llama-2-7B-32K-Instruct is an open-source, long-context chat model finetuned from Llama-2-7B-32K, over high-quality instruction and chat data. We built Llama-2-7B-32K-Instruct with less than 200 lines of Python script using Together API, and we also make the recipe fully availabl…

Open Source 7.0B ↓ 687
🤖
Qwen3-Swallow-30B-A3B-SFT-v0.2
tokyotech-llm

Qwen3-Swallow v0.2 is a family of large language models available in 8B , 30B-A3B , and 32B parameter sizes. Built as bilingual Japanese-English models, they were developed through Continual Pre-Training (CPT), Supervised Fine-Tuning (SFT), and Reinforcement Learning with Verifia…

Open Source 30.0B ↓ 682
🤖
DeepSeek-Math-V2
deepseek-ai

DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning

Open Source ↓ 670
🤖
Falcon3-3B-Instruct-1.58bit
tiiuae

0. TL;DR 1. Model Details 2. Training Details 3. Usage 4. Evaluation 5. Citation

Open Source 3.0B ↓ 667
🤖
Baichuan2-13B-Base
baichuan-inc

🦉GitHub 💬WeChat 百川API支持搜索增强和192K长窗口,新增百川搜索增强知识库、限时免费! 🚀 百川大模型在线对话平台 已正式向公众开放 🎉

Open Source 13.0B ↓ 656
🤖
Llama-3.1-Swallow-8B-Instruct-v0.1
tokyotech-llm

Llama 3.1 Swallow is a series of large language models (8B, 70B) that were built by continual pre-training on the Meta Llama 3.1 models. Llama 3.1 Swallow enhanced the Japanese language capabilities of the original Llama 3.1 while retaining the English language capabilities. We u…

Open Source 8.0B ↓ 652