AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,877 models Compare
🤖
Llama-SEA-LION-v3.5-70B-R-NVFP4
aisingapore

SEA-LION is a collection of Large Language Models (LLMs) which have been pretrained and instruct-tuned for the Southeast Asia (SEA) region.

Open Source 70.0B ↓ 143
🤖
internlm2-math-base-7b
internlm

State-of-the-art bilingual open-sourced Math reasoning LLMs. A solver , prover , verifier , augmentor .

Open Source 7.0B ↓ 142
🤖
CapRL-Eval-3B
internlm

CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning

Open Source 3.0B ↓ 142
🤖
GPT-OSS-Swallow-120B-SFT-v0.1
tokyotech-llm

GPT-OSS-Swallow v0.1 is a family of large language models available in 20B and 120B parameter sizes. Built as bilingual Japanese-English models, they were developed through Continual Pre-Training (CPT), Supervised Fine-Tuning (SFT), and Reinforcement Learning with Verifiable Rewa…

Open Source 120.0B ↓ 140
🤖
RLVR-8B-0926
stepfun-ai

PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning

Open Source 8.0B ↓ 139
🤖
SimpleSD-4B-thinking
apple

This model is an example of the Simple Self-Distillation (SimpleSD) method that improves code generation by fine-tuning a language model on its own sampled outputs—without rewards, verifiers, teacher models, or reinforcement learning. Please see the paper below for more informati…

Open Source 4.0B ↓ 138
🤖
bloom-7b1-intermediate
bigscience

WARNING: The checkpoints on this repo are not fully trained model. Evaluations of intermediary checkpoints and the final model will be added when conducted (see below).

Open Source ↓ 137
🤖
c4ai-command-r-plus
CohereLabs

Open Source ↓ 137
🤖
LongCat-Flash-Thinking
meituan-longcat

We introduce and release LongCat-Flash-Thinking , which is a powerful and efficient large reasoning model (LRM) with 560 billion total parameters, featuring an innovative Mixture-of-Experts (MoE) architecture. The model incorporates a dynamic computation mechanism that activates…

Open Source ↓ 134
🤖
SuperApriel-15B-Instruct
ServiceNow-AI

A 15B-parameter token-mixer supernet with 8 optimized deployment presets spanning 1.0× to 10.7× decode throughput at 32K sequence length, all from a single checkpoint. Derived from Apriel-1.6 through stochastic distillation and targeted supervised fine-tuning.

Open Source 15.0B ↓ 133
🤖
GPT-JT-Moderation-6B
togethercomputer

This model card introduces a moderation model, a GPT-JT model fine-tuned on Ontocord.ai's OIG-moderation dataset v0.1.

Open Source 6.0B ↓ 133
🤖
Swallow-13b-instruct-hf
tokyotech-llm

Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.

Open Source 13.0B ↓ 133