AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,877 models Compare
🤖
bloom-560m-intermediate
bigscience

WARNING: The checkpoints on this repo are not fully trained model. Evaluations of intermediary checkpoints and the final model will be added when conducted (see below).

Open Source ↓ 201
🤖
japanese-stablelm-base-alpha-7b
stabilityai

"A parrot able to speak Japanese, ukiyoe, edo period" — Stable Diffusion XL

Open Source 7.0B ↓ 200
🤖
AquilaMoE-SFT
BAAI

AquilaMoE: Efficient Training for MoE Models with Scale-Up and Scale-Out Strategies Language Foundation Model & Software Team Beijing Academy of Artificial Intelligence (BAAI) [Paper(released soon)] [Code] [github]

Open Source ↓ 199
🤖
japanese-stablelm-3b-4e1t-base
stabilityai

This is a 3B-parameter decoder-only language model with a focus on maximizing Japanese language modeling performance and Japanese downstream task performance. We conducted continued pretraining using Japanese data on the English language model, StableLM-3B-4E1T, to transfer the m…

Open Source 3.0B ↓ 197
🤖
internlm2-math-20b
internlm

State-of-the-art bilingual open-sourced Math reasoning LLMs. A solver , prover , verifier , augmentor .

Open Source 20.0B ↓ 195
🤖
AI21-Jamba2-Mini-FP8
ai21labs

Jamba2 Mini is an open source small language model built for enterprise reliability. With 12B active parameters (52B total), it delivers precise question answering without the computational overhead of reasoning models. The model's SSM-Transformer architecture provides a memory-e…

Open Source ↓ 193
🤖
Gemma-SEA-Guard-12B-2602
aisingapore

SEA-Guard is a collection of safety-focused Large Language Models (LLMs) designed specifically for the Southeast Asia (SEA) region. While the collection comprises four distinct models, we currently offer a single API endpoint that exclusively serves the Gemma-based model. You can…

Open Source 12.0B ↓ 192
🤖
Gemma-SEA-LION-v4-27B-VL
aisingapore

SEA-LION-VL is an instruct-tuned vision-text model for the Southeast Asia (SEA) region.

Multimodal 27.0B ↓ 191
🤖
starcoder-cxso
bigcode

An ablation of OctoCoder released for research purposes. Generally use OctoCoder, which performs better. Steps: 30

Code ↓ 189
🤖
Spatial-SSRL-3B
internlm

📖 Paper 🏠 Github 🤗 Spatial-SSRL-7B Model 🤗 Spatial-SSRL-3B Model 🤗 Spatial-SSRL-Qwen3VL-4B Model 🤗 Spatial-SSRL-81k Dataset 📰 Daily Paper

Open Source 3.0B ↓ 189
🤖
OREAL-7B
internlm

- Arxiv - Github - Model Collection - Data

Open Source 7.0B ↓ 188
🤖
Medical-Qwen3-Swallow-32B
tokyotech-llm

Medical-Qwen3-Swallow-32B is a medical-domain language model based on tokyotech-llm/Qwen3-Swallow-32B-RL-v0.2. It is designed to support research and development toward safe and trustworthy AI for Japanese clinical settings.

Open Source 32.0B ↓ 184