AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

440 models for "Base Model" Compare
Gemma-2-Llama-Swallow-27b-it-v0.1 logo
Gemma-2-Llama-Swallow-27b-it-v0.1
tokyotech-llm

Gemma-2-Llama-Swallow series was built by continual pre-training on the gemma-2 models. Gemma 2 Swallow enhanced the Japanese language capabilities of the original Gemma 2 while retaining the English language capabilities. We use approximately 200 billion tokens that were sampled…

Open Source 27.0B ↓ 261
Spatial-SSRL-Qwen3VL-4B logo
Spatial-SSRL-Qwen3VL-4B
internlm

📖 Paper 🏠 Github 🤗 Spatial-SSRL-7B Model 🤗 Spatial-SSRL-3B Model 🤗 Spatial-SSRL-Qwen3VL-4B Model 🤗 Spatial-SSRL-81k Dataset 📰 Daily Paper

Multimodal 4.0B ↓ 259
Youtu-Parsing logo
Youtu-Parsing
tencent

📃 License • 👨‍💻 Code • 🖥️ Demo • 📑 Technical Report • 📊 Benchmarks • 🚀 Getting Started

Open Source ↓ 255
AquilaMoE-SFT logo
AquilaMoE-SFT
BAAI

AquilaMoE: Efficient Training for MoE Models with Scale-Up and Scale-Out Strategies Language Foundation Model & Software Team Beijing Academy of Artificial Intelligence (BAAI) [Paper(released soon)] [Code] [github]

Open Source ↓ 254
Yi-34B-Chat-4bits logo
Yi-34B-Chat-4bits
01-ai

Building the Next Generation of Open-Source and Bilingual LLMs

Open Source 34.0B ↓ 254
BFS-Prover-V1-7B logo
BFS-Prover-V1-7B
ByteDance-Seed

🚀 BFS-Prover: Scalable Best-First Tree Search for LLM-based Automatic Theorem Proving State-of-the-art tactic generation model in Lean4

Open Source 7.0B ↓ 248
Baichuan-M2-32B-GPTQ-Int4 logo
Baichuan-M2-32B-GPTQ-Int4
baichuan-inc

Baichuan-M2-32B is Baichuan AI's medical-enhanced reasoning model, the second medical model released by Baichuan. Designed for real-world medical reasoning tasks, this model builds upon Qwen2.5-32B with an innovative Large Verifier System. Through domain-specific fine-tuning on r…

Open Source 32.0B ↓ 243
SimpleSD-4B-thinking logo
SimpleSD-4B-thinking
apple

This model is an example of the Simple Self-Distillation (SimpleSD) method that improves code generation by fine-tuning a language model on its own sampled outputs—without rewards, verifiers, teacher models, or reinforcement learning. Please see the paper below for more informati…

Open Source 4.0B ↓ 239
internlm2-chat-20b-sft logo
internlm2-chat-20b-sft
internlm

💻Github Repo • 🤔Reporting Issues • 📜Technical Report

Open Source 20.0B ↓ 234
Swallow-MX-8x7b-NVE-v0.1 logo
Swallow-MX-8x7b-NVE-v0.1
tokyotech-llm

Our Swallow-MX-8x7b-NVE-v0.1 model has undergone continuous pre-training from the Mixtral-8x7B-Instruct-v0.1, primarily with the addition of Japanese language data.

Open Source 7.0B ↓ 233
AquilaDense-7B logo
AquilaDense-7B
BAAI

AquilaMoE: Efficient Training for MoE Models with Scale-Up and Scale-Out Strategies Language Foundation Model & Software Team Beijing Academy of Artificial Intelligence (BAAI) [Paper(released soon)] [Code] [github]

Open Source 7.0B ↓ 229
SEA-LION-v1-3B logo
SEA-LION-v1-3B
aisingapore

SEA-LION is a collection of Large Language Models (LLMs) which has been pretrained and instruct-tuned for the Southeast Asia (SEA) region. The size of the models range from 3 billion to 7 billion parameters. This is the card for SEA-LION-v1-3B.

Open Source 3.0B ↓ 226