AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,302 models for "Transformer" Compare
🤖
Nous-Hermes-13b
NousResearch

Nous-Hermes-13b is a state-of-the-art language model fine-tuned on over 300,000 instructions. This model was fine-tuned by Nous Research, with Teknium and Karan4D leading the fine tuning process and dataset curation, Redmond AI sponsoring the compute, and several other contributo…

Open Source 13.0B ↓ 402
🤖
xLAM-7b-fc-r
Salesforce

[Homepage] [Paper] [Discord] [Dataset] [Github]

Open Source 7.0B ↓ 401
🤖
Seed-Coder-8B-Reasoning
ByteDance-Seed

Introduction We are thrilled to introduce Seed-Coder, a powerful, transparent, and parameter-efficient family of open-source code models at the 8B scale, featuring base, instruct, and reasoning variants. Seed-Coder contributes to promote the evolution of open code models through…

Code 8.0B ↓ 400
🤖
Stable-DiffCoder-8B-Instruct
ByteDance-Seed

Introduction We are thrilled to introduce Stable-DiffCoder, which is a strong code diffusion large language model. Built directly on the Seed-Coder architecture, data, and training pipeline, it introduces a block diffusion continual pretraining (CPT) stage with a tailored warmup…

Code 8.0B ↓ 398
🤖
Llama-SEA-LION-v3-70B-IT
aisingapore

SEA-LION is a collection of Large Language Models (LLMs) which have been pretrained and instruct-tuned for the Southeast Asia (SEA) region.

Open Source 70.0B ↓ 397
🤖
Yi-34B-Chat-8bits
01-ai

Building the Next Generation of Open-Source and Bilingual LLMs

Open Source 34.0B ↓ 397
🤖
Hunyuan-0.5B-Instruct
tencent

🤗  HuggingFace     🤖  ModelScope     🪡  AngelSlim

Open Source 0.5B ↓ 392
🤖
nomos-1
NousResearch

We release Nomos 1 , a specialization of Qwen/Qwen3-30B-A3B-Thinking-2507 for mathematical problem-solving and proof-writing in natural language. Nomos-1 was trained in collaboration with Hillclimb AI.

Open Source ↓ 391
🤖
falcon-rw-7b
tiiuae

Falcon-RW-7B is a 7B parameters causal decoder-only model built by TII and trained on 350B tokens of RefinedWeb. It is made available under the Apache 2.0 license.

Open Source 7.0B ↓ 390
🤖
internlm-20b
internlm

The Shanghai Artificial Intelligence Laboratory, in collaboration with SenseTime Technology, the Chinese University of Hong Kong, and Fudan University, has officially released the 20 billion parameter pretrained model, InternLM-20B. InternLM-20B was pre-trained on over 2.3T Token…

Open Source 20.0B ↓ 386
🤖
japanese-gpt-neox-3.6b-instruction-ppo
rinna

Overview This repository provides a Japanese GPT-NeoX model of 3.6 billion parameters. The model is based on rinna/japanese-gpt-neox-3.6b-instruction-sft-v2 and has been aligned to serve as an instruction-following conversational agent.

Open Source 3.6B ↓ 381
🤖
octocoder
bigcode

1. Model Summary 2. Use 3. Training 4. Citation

Code ↓ 379