AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,877 models Compare
🤖
NOSA-8B
openbmb

NOSA: Native and Offloadable Sparse Attention

Open Source 8.0B ↓ 306
🤖
Intern-S1-FP8
internlm

💻Github Repo • 🤗Model Collections • 📜Technical Report • 💬Online Chat

Open Source ↓ 304
🤖
internlm3-8b-instruct-awq
internlm

💻Github Repo • 🤔Reporting Issues • 📜Technical Report

Open Source 8.0B ↓ 304
🤖
MiMo-7B-RL-0530
XiaomiMiMo

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ Unlocking the Reasoning Potential of Language Model From Pretraining to Posttraining ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

Open Source 7.0B ↓ 302
🤖
aya-23-35B
CohereLabs

Open Source 35.0B ↓ 302
🤖
Llama-3.1-Swallow-8B-v0.1
tokyotech-llm

Llama 3.1 Swallow is a series of large language models (8B, 70B) that were built by continual pre-training on the Meta Llama 3.1 models. Llama 3.1 Swallow enhanced the Japanese language capabilities of the original Llama 3.1 while retaining the English language capabilities. We u…

Open Source 8.0B ↓ 301
🤖
UI-TARS-72B-DPO
ByteDance-Seed

UI-TARS-72B-DPO UI-TARS-2B-SFT     UI-TARS-7B-SFT     UI-TARS-7B-DPO (Recommended)     UI-TARS-72B-SFT     UI-TARS-72B-DPO (Recommended) Introduction

Open Source 72.0B ↓ 299
🤖
Penguin-VL-8B
tencent

Penguin-VL Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders

Multimodal 8.0B ↓ 298
🤖
DeepSeek-R1-Distill-Qwen-32B-Japanese
cyberagent

This is a Japanese finetuned model based on deepseek-ai/DeepSeek-R1-Distill-Qwen-32B.

Reasoning 32.0B ↓ 298
🤖
CoDA-v0-Instruct
Salesforce

CoDA: Coding LM via Diffusion Adaptation

Open Source ↓ 297
🤖
ArmorOCR
inclusionAI

ArmorOCR is a two-stage framework for grounded adversarial OCR perception built on Qwen3-VL-8B-Instruct. It enables single-pass inference on the original image, without any inference-time visual transformations or tool assistance.

Open Source ↓ 296
🤖
Llama-SEA-LION-v2-8B
aisingapore

Llama-SEA-LION-v2-8B SEA-LION is a collection of Large Language Models (LLMs) which have been pretrained and instruct-tuned for the Southeast Asia (SEA) region.

Open Source 8.0B ↓ 293