AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,877 models Compare
🤖
Nous-Hermes-2-Vision-Alpha
NousResearch

In the tapestry of Greek mythology, Hermes reigns as the eloquent Messenger of the Gods, a deity who deftly bridges the realms through the art of communication. It is in homage to this divine mediator that I name this advanced LLM "Hermes," a system crafted to navigate the comple…

Multimodal ↓ 225
🤖
japanese-stablelm-base-beta-70b
stabilityai

A cute robot wearing a kimono writes calligraphy with one single brush — Stable Diffusion XL

Open Source 70.0B ↓ 225
🤖
CADD-Base-7B
apple

CADD-Base-7B is a masked diffusion language model for code generation, augmented with Continuously Augmented Discrete Diffusion (CADD) --- a continuous flow-matching signal that guides the discrete denoising process.

Open Source 7.0B ↓ 225
🤖
SmolLM3-3B-ONNX
HuggingFaceTB

1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. License

Open Source 3.0B ↓ 224
🤖
SEA-LION-v1-7B-IT-Research
aisingapore

SEA-LION is a collection of Large Language Models (LLMs) which has been pretrained and instruct-tuned for the Southeast Asia (SEA) region. The size of the models range from 3 billion to 7 billion parameters. This is the card for the SEA-LION 7B Instruct (Non-Commercial) model.

Open Source 7.0B ↓ 223
🤖
codegen-2B-nl
Salesforce

CodeGen is a family of autoregressive language models for program synthesis from the paper: A Conversational Paradigm for Program Synthesis by Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, Caiming Xiong. The models are originally releas…

Code 2.0B ↓ 222
🤖
CapRL-3B
internlm

CapRL 📖 Paper 🏠 Github 🤗 CapRL Collection 🤗 Daily Paper

Open Source 3.0B ↓ 220
🤖
AquilaCode-py
BAAI

Aquila Language Model is the first open source language model that supports both Chinese and English knowledge, commercial license agreements, and compliance with domestic data regulations.

Code ↓ 220
🤖
Llama-3.3-Swallow-70B-v0.4
tokyotech-llm

Llama 3.3 Swallow is a large language model (70B) that was built by continual pre-training on the Meta Llama 3.3 model. Llama 3.3 Swallow enhanced the Japanese language capabilities of the original Llama 3.3 while retaining the English language capabilities. We use approximately…

Open Source 70.0B ↓ 215
🤖
SynLogic-Mix-3-32B
MiniMaxAI

SynLogic Zero-Mix-3: Large-Scale Multi-Domain Reasoning Model

Open Source 32.0B ↓ 215
🤖
stablelm-base-alpha-3b-v2
stabilityai

StableLM-Base-Alpha-3B-v2 is a 3 billion parameter decoder-only language model pre-trained on diverse English datasets. This model is the successor to the first StableLM-Base-Alpha-3B model, addressing previous shortcomings through the use of improved data sources and mixture rat…

Open Source 3.0B ↓ 214
🤖
starcoderbase-7b
bigcode

Code 7.0B ↓ 213