AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

440 models for "Base Model" Compare
CapRL-3B logo
CapRL-3B
internlm

CapRL 📖 Paper 🏠 Github 🤗 CapRL Collection 🤗 Daily Paper

Open Source 3.0B ↓ 370
Gemma-2-Llama-Swallow-2b-it-v0.1 logo
Gemma-2-Llama-Swallow-2b-it-v0.1
tokyotech-llm

Gemma-2-Llama-Swallow series was built by continual pre-training on the gemma-2 models. Gemma 2 Swallow enhanced the Japanese language capabilities of the original Gemma 2 while retaining the English language capabilities. We use approximately 200 billion tokens that were sampled…

Open Source 2.0B ↓ 367
AquilaMoE logo
AquilaMoE
BAAI

AquilaMoE: Efficient Training for MoE Models with Scale-Up and Scale-Out Strategies Language Foundation Model & Software Team Beijing Academy of Artificial Intelligence (BAAI) [Paper(released soon)] [Code] [github]

Open Source ↓ 365
ArmorOCR logo
ArmorOCR
inclusionAI

ArmorOCR is a two-stage framework for grounded adversarial OCR perception built on Qwen3-VL-8B-Instruct. It enables single-pass inference on the original image, without any inference-time visual transformations or tool assistance.

Open Source ↓ 361
Yi-6B-Chat-4bits logo
Yi-6B-Chat-4bits
01-ai

Building the Next Generation of Open-Source and Bilingual LLMs

Open Source 6.0B ↓ 356
CapRL-InternVL3.5-8B logo
CapRL-InternVL3.5-8B
internlm

CapRL 📖 Paper 🏠 Github 🤗 CapRL Collection 🤗 Daily Paper

Multimodal 8.0B ↓ 351
Spatial-SSRL-3B logo
Spatial-SSRL-3B
internlm

📖 Paper 🏠 Github 🤗 Spatial-SSRL-7B Model 🤗 Spatial-SSRL-3B Model 🤗 Spatial-SSRL-Qwen3VL-4B Model 🤗 Spatial-SSRL-81k Dataset 📰 Daily Paper

Open Source 3.0B ↓ 349
Baichuan-M3-235B-GPTQ-INT4 logo
Baichuan-M3-235B-GPTQ-INT4
baichuan-inc

From Inquiry to Decision: Building Trustworthy Medical AI

Open Source 235.0B ↓ 349
Llama-3.1-Swallow-70B-Instruct-v0.3 logo
Llama-3.1-Swallow-70B-Instruct-v0.3
tokyotech-llm

Llama 3.1 Swallow is a series of large language models (8B, 70B) that were built by continual pre-training on the Meta Llama 3.1 models. Llama 3.1 Swallow enhanced the Japanese language capabilities of the original Llama 3.1 while retaining the English language capabilities. We u…

Open Source 70.0B ↓ 344
Ling-2.6-flash-base logo
Ling-2.6-flash-base
inclusionAI

🤗 Hugging Face      🤖 ModelScope       Tech Report       💻 GitHub

Open Source ↓ 344
SynLogic-32B logo
SynLogic-32B
MiniMaxAI

SynLogic-32B: Advanced Logical Reasoning Model

Open Source 32.0B ↓ 339
POINTS-GUI-G logo
POINTS-GUI-G
tencent

- 🔜 Upcoming: The End-to-End GUI Agent Model is currently under active development and will be released in a subsequent update. Stay tuned! - 🚀 2026.02.06: We are happy to present POINTS-GUI-G , our specialized GUI Grounding Model. To facilitate reproducible evaluation, we provid…

Open Source ↓ 339