AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,299 models for "Transformer" Compare
🤖
Step-3.5-Flash-Base
stepfun-ai

Step 3.5 Flash (visit website) is our most capable open-source foundation model, engineered to deliver frontier reasoning and agentic capabilities with exceptional efficiency. We also open-sourced the training codebase (SteptronOss), with support for continue pretrain, SFT, RL (W…

Open Source ↓ 321
🤖
codegen25-7b-multi_P
Salesforce

Authors: Erik Nijkamp\ , Hiroaki Hayashi\ , Yingbo Zhou, Caiming Xiong

Code 7.0B ↓ 319
🤖
Medical-Qwen3-Swallow-30B-A3B
tokyotech-llm

Medical-Qwen3-Swallow-30B-A3B is a medical-domain language model based on tokyotech-llm/Qwen3-Swallow-30B-A3B-RL-v0.2 . It is designed to support research and development toward safe and trustworthy AI for Japanese clinical settings.

Open Source 30.0B ↓ 310
🤖
Llama-2-70b-instruct
upstage

Developed by : Upstage Backbone Model : LLaMA-2 Language(s) : English Library : HuggingFace Transformers License : Fine-tuned checkpoints is licensed under the Non-Commercial Creative Commons license (CC BY-NC-4.0) Where to send comments : Instructions on how to provide feedback…

Open Source 70.0B ↓ 307
🤖
Ring-mini-2.0
inclusionAI

🤗 Hugging Face &nbsp&nbsp &nbsp&nbsp🤖 ModelScope   🐙 Experience Now

Open Source ↓ 306
🤖
Intern-S1-FP8
internlm

💻Github Repo • 🤗Model Collections • 📜Technical Report • 💬Online Chat

Open Source ↓ 304
🤖
internlm3-8b-instruct-awq
internlm

💻Github Repo • 🤔Reporting Issues • 📜Technical Report

Open Source 8.0B ↓ 304
🤖
MiMo-7B-RL-0530
XiaomiMiMo

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ Unlocking the Reasoning Potential of Language Model From Pretraining to Posttraining ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

Open Source 7.0B ↓ 302
🤖
Penguin-VL-8B
tencent

Penguin-VL Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders

Multimodal 8.0B ↓ 298
🤖
DeepSeek-R1-Distill-Qwen-32B-Japanese
cyberagent

This is a Japanese finetuned model based on deepseek-ai/DeepSeek-R1-Distill-Qwen-32B.

Reasoning 32.0B ↓ 298
🤖
CoDA-v0-Instruct
Salesforce

CoDA: Coding LM via Diffusion Adaptation

Open Source ↓ 297
🤖
ArmorOCR
inclusionAI

ArmorOCR is a two-stage framework for grounded adversarial OCR perception built on Qwen3-VL-8B-Instruct. It enables single-pass inference on the original image, without any inference-time visual transformations or tool assistance.

Open Source ↓ 296