AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,877 models Compare
🤖
CodeLlama-34b-Instruct-hf
meta-llama

Code 34.0B ↓ 1K
🤖
Qwen3-Swallow-8B-RL-v0.2-AWQ-INT4
tokyotech-llm

Qwen3-Swallow v0.2 is a family of large language models available in 8B , 30B-A3B , and 32B parameter sizes. Built as bilingual Japanese-English models, they were developed through Continual Pre-Training (CPT), Supervised Fine-Tuning (SFT), and Reinforcement Learning with Verifia…

Open Source 8.0B ↓ 1K
🤖
Nous-Hermes-2-Mistral-7B-DPO
NousResearch

Nous Hermes 2 on Mistral 7B DPO is the new flagship 7B Hermes! This model was DPO'd from Teknium/OpenHermes-2.5-Mistral-7B and has improved across the board on all benchmarks tested - AGIEval, BigBench Reasoning, GPT4All, and TruthfulQA.

Open Source 7.0B ↓ 1K
🤖
RedPajama-INCITE-Instruct-3B-v1
togethercomputer

RedPajama-INCITE-Instruct-3B-v1 was developed by Together and leaders from the open-source AI community including Ontocord.ai, ETH DS3Lab, AAI CERC, Université de Montréal, MILA - Québec AI Institute, Stanford Center for Research on Foundation Models (CRFM), Stanford Hazy Researc…

Open Source 3.0B ↓ 988
🤖
codegen-2B-multi
Salesforce

CodeGen is a family of autoregressive language models for program synthesis from the paper: A Conversational Paradigm for Program Synthesis by Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, Caiming Xiong. The models are originally releas…

Code 2.0B ↓ 985
🤖
EXAONE-Deep-7.8B
LGAI-EXAONE

We introduce EXAONE Deep, which exhibits superior capabilities in various reasoning tasks including math and coding benchmarks, ranging from 2.4B to 32B parameters developed and released by LG AI Research. Evaluation results show that 1) EXAONE Deep 2.4B outperforms other models…

Open Source 7.8B ↓ 984
🤖
StepFun-Formalizer-7B
stepfun-ai

StepFun-Formalizer: Unlocking the Autoformalization Potential of LLMs through Knowledge-Reasoning Fusion

Open Source 7.0B ↓ 984
🤖
GPT-OSS-Swallow-20B-SFT-v0.1
tokyotech-llm

GPT-OSS-Swallow v0.1 is a family of large language models available in 20B and 120B parameter sizes. Built as bilingual Japanese-English models, they were developed through Continual Pre-Training (CPT), Supervised Fine-Tuning (SFT), and Reinforcement Learning with Verifiable Rewa…

Open Source 20.0B ↓ 977
🤖
blip2-opt-6.7b-coco
Salesforce

BLIP-2 model, leveraging OPT-6.7b (a large language model with 6.7 billion parameters). It was introduced in the paper BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models by Li et al. and first released in this repository.

Open Source 6.7B ↓ 968
🤖
CodeLlama-7b-hf
meta-llama

Code 7.0B ↓ 957
🤖
grok-1
xai-org

This repository contains the weights of the Grok-1 open-weights model. You can find the code in the GitHub Repository.

Open Source ↓ 947
🤖
nanowhale-100m
HuggingFaceTB

A small ~110M parameter language model implementing the DeepSeek-V4 architecture , fine-tuned for chat/instruction following. Trained from scratch — no weights from DeepSeek-V4 were used.

Open Source ↓ 939