AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

440 models for "Base Model" Compare
xgen-7b-8k-inst logo
xgen-7b-8k-inst
Salesforce

Official research release for the family of XGen models ( 7B ) by Salesforce AI Research:

Open Source 7.0B ↓ 454
Ling-2.6-1T-base logo
Ling-2.6-1T-base
inclusionAI

🤗 Hugging Face      🤖 ModelScope       Tech Report       🐙 GitHub

Open Source ↓ 452
Aquila2-34B logo
Aquila2-34B
BAAI

We opensource our Aquila2 series, now including Aquila2 , the base language models, namely Aquila2-7B and Aquila2-34B , as well as AquilaChat2 , the chat models, namely AquilaChat2-7B and AquilaChat2-34B , as well as the long-text chat models, namely AquilaChat2-7B-16k and Aquila…

Open Source 34.0B ↓ 441
japanese-stablelm-instruct-gamma-7b logo
japanese-stablelm-instruct-gamma-7b
stabilityai

This is a 7B-parameter decoder-only Japanese language model fine-tuned on instruction-following datasets, built on top of the base model Japanese Stable LM Base Gamma 7B.

Open Source 7.0B ↓ 440
Swallow-7b-hf logo
Swallow-7b-hf
tokyotech-llm

Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.

Open Source 7.0B ↓ 433
Llama-3.3-Swallow-70B-v0.4 logo
Llama-3.3-Swallow-70B-v0.4
tokyotech-llm

Llama 3.3 Swallow is a large language model (70B) that was built by continual pre-training on the Meta Llama 3.3 model. Llama 3.3 Swallow enhanced the Japanese language capabilities of the original Llama 3.3 while retaining the English language capabilities. We use approximately…

Open Source 70.0B ↓ 419
internlm2_5-20b logo
internlm2_5-20b
internlm

💻Github Repo • 🤔Reporting Issues • 📜Technical Report

Open Source 20.0B ↓ 410
CapRL-Qwen3VL-2B logo
CapRL-Qwen3VL-2B
internlm

CapRL 📖 Paper 🏠 Github 🤗 CapRL Collection 🤗 Daily Paper

Multimodal 2.0B ↓ 399
japanese-stablelm-3b-4e1t-instruct logo
japanese-stablelm-3b-4e1t-instruct
stabilityai

This is a 3B-parameter decoder-only Japanese language model fine-tuned on instruction-following datasets, built on top of the base model Japanese StableLM-3B-4E1T Base.

Open Source 3.0B ↓ 389
CodeLlama-7b-hf logo
CodeLlama-7b-hf
NousResearch

Code 7.0B ↓ 383
Swallow-13b-instruct-hf logo
Swallow-13b-instruct-hf
tokyotech-llm

Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.

Open Source 13.0B ↓ 380
SmolLM3-3B-GSM8K-SFT logo
SmolLM3-3B-GSM8K-SFT
HuggingFaceTB

Fine-tuned version of HuggingFaceTB/SmolLM3-3B-Base optimized for grade school math (GSM8K benchmark).

Open Source 3.0B ↓ 371