AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

440 models for "Base Model" Compare
Apriel-5B-Base logo
Apriel-5B-Base
ServiceNow-AI

1. Model Summary 2. Evaluation 3. Intended Use 4. Limitations 5. Security and Responsible Use 6. License 7. Citation

Open Source 5.0B ↓ 329
Penguin-VL-2B logo
Penguin-VL-2B
tencent

Penguin-VL Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders

Multimodal 2.0B ↓ 329
SEA-LION-v1-7B logo
SEA-LION-v1-7B
aisingapore

SEA-LION is a collection of Large Language Models (LLMs) which has been pretrained and instruct-tuned for the Southeast Asia (SEA) region. The size of the models range from 3 billion to 7 billion parameters. This is the card for the SEA-LION 7B base model.

Open Source 7.0B ↓ 322
SimpleSD-4B-instruct logo
SimpleSD-4B-instruct
apple

This model is an example of the Simple Self-Distillation (SimpleSD) method that improves code generation by fine-tuning a language model on its own sampled outputs—without rewards, verifiers, teacher models, or reinforcement learning. Please see the paper below for more informati…

Open Source 4.0B ↓ 315
ZwZ-4B logo
ZwZ-4B
inclusionAI

ZwZ-4B is a fine-grained multimodal perception model built upon Qwen3-VL-4B. It is trained using Region-to-Image Distillation (R2I) combined with reinforcement learning, enabling superior fine-grained visual understanding in a single forward pass — no inference-time zooming or to…

Open Source 4.0B ↓ 309
Yarn-Llama-2-13b-64k logo
Yarn-Llama-2-13b-64k
NousResearch

Nous-Yarn-Llama-2-13b-64k is a state-of-the-art language model for long context, further pretrained on long context data for 400 steps. This model is the Flash Attention 2 patched version of the original model: https://huggingface.co/conceptofmind/Yarn-Llama-2-13b-64k

Open Source 13.0B ↓ 307
Llama-3.1-Swallow-70B-Instruct-v0.1 logo
Llama-3.1-Swallow-70B-Instruct-v0.1
tokyotech-llm

Llama 3.1 Swallow is a series of large language models (8B, 70B) that were built by continual pre-training on the Meta Llama 3.1 models. Llama 3.1 Swallow enhanced the Japanese language capabilities of the original Llama 3.1 while retaining the English language capabilities. We u…

Open Source 70.0B ↓ 307
Falcon3-10B-Base-1.58bit-prequantized logo
Falcon3-10B-Base-1.58bit-prequantized
tiiuae

--- library name: transformers tags: - bitnet - falcon3 base model: tiiuae/Falcon3-10B-Base license: other license name: falcon-llm-license license link: https://falconllm.tii.ae/falcon-terms-and-conditions.html ---

Open Source 10.0B ↓ 305
AgentCPM-Report logo
AgentCPM-Report
openbmb

AgentCPM-Report: Gemini-2.5-pro-DeepResearch Level Local DeepResearch

Open Source ↓ 302
Penguin-VL-8B logo
Penguin-VL-8B
tencent

Penguin-VL Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders

Multimodal 8.0B ↓ 298
Spatial-SSRL-7B logo
Spatial-SSRL-7B
internlm

📖 Paper 🏠 Github 🤗 Spatial-SSRL-7B Model 🤗 Spatial-SSRL-3B Model 🤗 Spatial-SSRL-Qwen3VL-4B Model 🤗 Spatial-SSRL-81k Dataset 📰 Daily Paper

Open Source 7.0B ↓ 297
MiMo-7B-RL-Zero logo
MiMo-7B-RL-Zero
XiaomiMiMo

━━━━━━━━━━━━━━━━━━━━━━━━━ Unlocking the Reasoning Potential of Language Model From Pretraining to Posttraining ━━━━━━━━━━━━━━━━━━━━━━━━━

Open Source 7.0B ↓ 295