AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

440 models for "Base Model" Compare
ERNIE-4.5-0.3B-Base-PT logo
ERNIE-4.5-0.3B-Base-PT
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Open Source 0.36B ↓ 2K
InternVL3-2B-Instruct logo
InternVL3-2B-Instruct
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 2.09B ↓ 2K
InternVL3-8B-Instruct logo
InternVL3-8B-Instruct
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 8.1B ↓ 2K
InternVL3-78B-Instruct logo
InternVL3-78B-Instruct
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 78.4B ↓ 2K
AprielGuard logo
AprielGuard
ServiceNow-AI

1. Summary 2. Taxonomy 2. Evaluation 3. Training Details 4. How to Use 5. Intended Use 6. Limitations 7. License 8. Citation

Open Source 8.0B ↓ 1.9K
SimpleSD-30B-instruct logo
SimpleSD-30B-instruct
apple

This model is an example of the Simple Self-Distillation (SimpleSD) method that improves code generation by fine-tuning a language model on its own sampled outputs—without rewards, verifiers, teacher models, or reinforcement learning. Please see the paper below for more informati…

Open Source 30.0B ↓ 1.9K
Swallow-7b-instruct-hf logo
Swallow-7b-instruct-hf
tokyotech-llm

Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.

Open Source 7.0B ↓ 1.9K
GLM-Z1-9B-0414 logo
GLM-Z1-9B-0414
zai-org

The GLM family welcomes a new generation of open-source models, the GLM-4-32B-0414 series, featuring 32 billion parameters. Its performance is comparable to OpenAI's GPT series and DeepSeek's V3/R1 series, and it supports very user-friendly local deployment features. GLM-4-32B-Ba…

Open Source 9.0B ↓ 1.6K
AREX-2 logo
AREX-2
BAAI

AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks

Open Source 27.0B ↓ 1.5K
internlm2_5-20b-chat logo
internlm2_5-20b-chat
internlm

💻Github Repo • 🤔Reporting Issues • 📜Technical Report

Open Source 20.0B ↓ 1.5K
Youtu-LLM-2B-Base logo
Youtu-LLM-2B-Base
tencent

📃 License • 💻 Code • 📑 Technical Report • 📊 Benchmarks

Open Source 2.0B ↓ 1.5K
Hunyuan-0.5B-Instruct logo
Hunyuan-0.5B-Instruct
tencent

🤗  HuggingFace     🤖  ModelScope     🪡  AngelSlim

Open Source 0.5B ↓ 1.4K