AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,877 models Compare
🤖
japanese-stablelm-instruct-beta-7b
stabilityai

A cute robot wearing a kimono writes calligraphy with one single brush — Stable Diffusion XL

Open Source 7.0B ↓ 178
🤖
Qwen-SEA-LION-v4-32B-IT-4BIT
aisingapore

Qwen-SEA-LION-v4-32B-IT-4BIT (GPTQ model)

Open Source 32.0B ↓ 177
🤖
japanese-stablelm-instruct-alpha-7b-v2
stabilityai

"A parrot able to speak Japanese, ukiyoe, edo period" — Stable Diffusion XL

Open Source 7.0B ↓ 176
🤖
AquilaDense-16B
BAAI

AquilaMoE: Efficient Training for MoE Models with Scale-Up and Scale-Out Strategies Language Foundation Model & Software Team Beijing Academy of Artificial Intelligence (BAAI) [Paper(released soon)] [Code] [github]

Open Source 16.0B ↓ 176
🤖
Medical-GPT-OSS-Swallow-120B
tokyotech-llm

Medical-GPT-OSS-Swallow-120B is a medical-domain language model based on tokyotech-llm/GPT-OSS-Swallow-120B-RL-v0.1. It is designed to support research and development toward safe and trustworthy AI for Japanese clinical settings.

Open Source 120.0B ↓ 175
🤖
llama-65b-instruct
upstage

Developed by : Upstage Backbone Model : LLaMA Variations : It has different model parameter sizes and sequence lengths: 30B/1024, 30B/2048, 65B/1024 Language(s) : English Library : HuggingFace Transformers License : This model is under a Non-commercial Bespoke License and governe…

Open Source 65.0B ↓ 175
🤖
internlm2-chat-1_8b-sft
internlm

💻Github Repo • 🤔Reporting Issues • 📜Technical Report

Open Source 8.0B ↓ 173
🤖
octogeex
bigcode

1. Model Summary 2. Use 3. Training 4. License 5. Citation

Open Source ↓ 173
🤖
Ring-lite
inclusionAI

Ring-lite is a lightweight, fully open-sourced MoE (Mixture of Experts) LLM designed for complex reasoning tasks. It is built upon the publicly available Ling-lite-1.5 model, which has 16.8B parameters with 2.75B activated parameters.. We use a joint training pipeline combining k…

Open Source ↓ 172
🤖
SimpleSD-30B-instruct
apple

This model is an example of the Simple Self-Distillation (SimpleSD) method that improves code generation by fine-tuning a language model on its own sampled outputs—without rewards, verifiers, teacher models, or reinforcement learning. Please see the paper below for more informati…

Open Source 30.0B ↓ 172
🤖
AHN-Mamba2-for-Qwen-2.5-Instruct-14B
ByteDance-Seed

AHN: Artificial Hippocampus Networks for Efficient Long-Context Modeling

Open Source 14.0B ↓ 172
🤖
Step-3.5-Flash-Base-Midtrain
stepfun-ai

Step 3.5 Flash (visit website) is our most capable open-source foundation model, engineered to deliver frontier reasoning and agentic capabilities with exceptional efficiency. We also open-sourced the training codebase (SteptronOss), with support for continue pretrain, SFT, RL (W…

Open Source ↓ 168