AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,877 models Compare
🤖
nomos-1
NousResearch

We release Nomos 1 , a specialization of Qwen/Qwen3-30B-A3B-Thinking-2507 for mathematical problem-solving and proof-writing in natural language. Nomos-1 was trained in collaboration with Hillclimb AI.

Open Source ↓ 391
🤖
falcon-rw-7b
tiiuae

Falcon-RW-7B is a 7B parameters causal decoder-only model built by TII and trained on 350B tokens of RefinedWeb. It is made available under the Apache 2.0 license.

Open Source 7.0B ↓ 390
🤖
internlm-20b
internlm

The Shanghai Artificial Intelligence Laboratory, in collaboration with SenseTime Technology, the Chinese University of Hong Kong, and Fudan University, has officially released the 20 billion parameter pretrained model, InternLM-20B. InternLM-20B was pre-trained on over 2.3T Token…

Open Source 20.0B ↓ 386
🤖
japanese-gpt-neox-3.6b-instruction-ppo
rinna

Overview This repository provides a Japanese GPT-NeoX model of 3.6 billion parameters. The model is based on rinna/japanese-gpt-neox-3.6b-instruction-sft-v2 and has been aligned to serve as an instruction-following conversational agent.

Open Source 3.6B ↓ 381
🤖
octocoder
bigcode

1. Model Summary 2. Use 3. Training 4. Citation

Code ↓ 379
🤖
Apriel-1.5-15b-Thinker
ServiceNow-AI

Apriel-1.5-15b-Thinker - Mid training is all you need!

Open Source 15.0B ↓ 379
🤖
OpenELM-450M
apple

Sachin Mehta, Mohammad Hossein Sekhavat, Qingqing Cao, Maxwell Horton, Yanzi Jin, Chenfan Sun, Iman Mirzadeh, Mahyar Najibi, Dmitry Belenko, Peter Zatloukal, Mohammad Rastegari

Open Source ↓ 379
🤖
MiMo-7B-RL-Zero
XiaomiMiMo

━━━━━━━━━━━━━━━━━━━━━━━━━ Unlocking the Reasoning Potential of Language Model From Pretraining to Posttraining ━━━━━━━━━━━━━━━━━━━━━━━━━

Open Source 7.0B ↓ 378
🤖
Yarn-Llama-2-13b-128k
NousResearch

Nous-Yarn-Llama-2-13b-128k is a state-of-the-art language model for long context, further pretrained on long context data for 600 steps. This model is the Flash Attention 2 patched version of the original model: https://huggingface.co/conceptofmind/Yarn-Llama-2-13b-128k

Open Source 13.0B ↓ 378
🤖
japanese-stablelm-instruct-gamma-7b
stabilityai

This is a 7B-parameter decoder-only Japanese language model fine-tuned on instruction-following datasets, built on top of the base model Japanese Stable LM Base Gamma 7B.

Open Source 7.0B ↓ 377
🤖
tiny-aya-water
CohereLabs

Open Source ↓ 375
🤖
Yi-6B-Chat-8bits
01-ai

Building the Next Generation of Open-Source and Bilingual LLMs

Open Source 6.0B ↓ 373