AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

440 models for "Base Model" Compare
deepseek-coder-5.7bmqa-base logo
deepseek-coder-5.7bmqa-base
deepseek-ai

[🏠Homepage] [🤖 Chat with DeepSeek Coder] [Discord] [Wechat(微信)]

Code 5.7B ↓ 918
Llama-2-70b-hf logo
Llama-2-70b-hf
NousResearch

Llama 2 Llama 2 is a collection of pretrained and fine-tuned generative text models ranging in scale from 7 billion to 70 billion parameters. This is the repository for the 70B pretrained model, converted for the Hugging Face Transformers format. Links to other models can be foun…

Open Source 70.0B ↓ 899
ContextPilot-14B logo
ContextPilot-14B
tencent

ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL

Open Source 14.8B ↓ 897
RedPajama-INCITE-7B-Chat logo
RedPajama-INCITE-7B-Chat
togethercomputer

RedPajama-INCITE-7B-Chat was developed by Together and leaders from the open-source AI community including Ontocord.ai, ETH DS3Lab, AAI CERC, Université de Montréal, MILA - Québec AI Institute, Stanford Center for Research on Foundation Models (CRFM), Stanford Hazy Research resea…

Open Source 7.0B ↓ 885
Hunyuan-1.8B-Pretrain logo
Hunyuan-1.8B-Pretrain
tencent

🤗  HuggingFace     🤖  ModelScope     🪡  AngelSlim

Open Source 1.8B ↓ 883
RedPajama-INCITE-Instruct-3B-v1 logo
RedPajama-INCITE-Instruct-3B-v1
togethercomputer

RedPajama-INCITE-Instruct-3B-v1 was developed by Together and leaders from the open-source AI community including Ontocord.ai, ETH DS3Lab, AAI CERC, Université de Montréal, MILA - Québec AI Institute, Stanford Center for Research on Foundation Models (CRFM), Stanford Hazy Researc…

Open Source 3.0B ↓ 849
ERNIE-4.5-VL-424B-A47B-Base-PT logo
ERNIE-4.5-VL-424B-A47B-Base-PT
baidu

[!NOTE] Note: " -Paddle " models use PaddlePaddle weights, while " -PT " models use Transformer-style PyTorch weights.

Multimodal 424.0B ↓ 840
Medical-Qwen3-Swallow-8B logo
Medical-Qwen3-Swallow-8B
tokyotech-llm

Medical-Qwen3-Swallow-8B is a medical-domain language model based on tokyotech-llm/Qwen3-Swallow-8B-RL-v0.2. It is designed to support research and development toward safe and trustworthy AI for Japanese clinical settings.

Open Source 8.0B ↓ 824
Llama-3.1-Swallow-8B-Instruct-v0.1 logo
Llama-3.1-Swallow-8B-Instruct-v0.1
tokyotech-llm

Llama 3.1 Swallow is a series of large language models (8B, 70B) that were built by continual pre-training on the Meta Llama 3.1 models. Llama 3.1 Swallow enhanced the Japanese language capabilities of the original Llama 3.1 while retaining the English language capabilities. We u…

Open Source 8.0B ↓ 820
UI-Mate-9B logo
UI-Mate-9B
tencent

UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations

Open Source 9.0B ↓ 808
InternVL3-38B-AWQ logo
InternVL3-38B-AWQ
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 38.0B ↓ 792
Hunyuan-4B-Instruct logo
Hunyuan-4B-Instruct
tencent

🤗  HuggingFace     🤖  ModelScope     🪡  AngelSlim

Open Source 4.0B ↓ 770