AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,305 models for "Transformer" Compare
Falcon-H1-34B-Base logo
Falcon-H1-34B-Base
tiiuae

0. TL;DR 1. Model Details 2. Training Details 3. Usage 4. Evaluation 5. Citation

open-source 34.0B ↓ 633
MiniCPM-2B-dpo-bf16 logo
MiniCPM-2B-dpo-bf16
openbmb

MiniCPM 技术报告 Technical Report OmniLMM 多模态模型 Multi-modal Model CPM-C 千亿模型试用 ~100B Model Trial

open-source 2.0B ↓ 629
Seed-Coder-8B-Base logo
Seed-Coder-8B-Base
ByteDance-Seed

Introduction We are thrilled to introduce Seed-Coder, a powerful, transparent, and parameter-efficient family of open-source code models at the 8B scale, featuring base, instruct, and reasoning variants. Seed-Coder contributes to promote the evolution of open code models through…

code 8.0B ↓ 627
HY-Embodied-0.5-X logo
HY-Embodied-0.5-X
tencent

HY-Embodied-0.5-X An Enhanced Embodied Foundation Model for Real-World Agents Tencent Robotics X × HY Vision Team

open-source ↓ 625
LLaDA2.0-mini-preview logo
LLaDA2.0-mini-preview
inclusionAI

LLaDA2.0-mini-preview is a diffusion language model featuring a 16BA1B Mixture-of-Experts (MoE) architecture. As an enhanced, instruction-tuned iteration of the LLaDA series, it is optimized for practical applications.

open-source ↓ 618
Falcon-E-1B-Base logo
Falcon-E-1B-Base
tiiuae

0. TL;DR 1. Model Details 2. Training Details 3. Usage 4. Evaluation 5. Citation

open-source 1.0B ↓ 610
calm3-22b-chat logo
calm3-22b-chat
cyberagent

CyberAgentLM3 is a decoder-only language model pre-trained on 2.0 trillion tokens from scratch. CyberAgentLM3-Chat is a fine-tuned model specialized for dialogue use cases.

open-source 22.0B ↓ 607
RedPajama-INCITE-Chat-3B-v1 logo
RedPajama-INCITE-Chat-3B-v1
togethercomputer

RedPajama-INCITE-Chat-3B-v1 was developed by Together and leaders from the open-source AI community including Ontocord.ai, ETH DS3Lab, AAI CERC, Université de Montréal, MILA - Québec AI Institute, Stanford Center for Research on Foundation Models (CRFM), Stanford Hazy Research re…

open-source 3.0B ↓ 601
Falcon-H1-1.5B-Deep-Base logo
Falcon-H1-1.5B-Deep-Base
tiiuae

0. TL;DR 1. Model Details 2. Training Details 3. Usage 4. Evaluation 5. Citation

open-source 1.5B ↓ 600
DeepHermes-3-Llama-3-3B-Preview logo
DeepHermes-3-Llama-3-3B-Preview
NousResearch

DeepHermes 3 Preview is the latest version of our flagship Hermes series of LLMs by Nous Research, and one of the first models in the world to unify Reasoning (long chains of thought that improve answer accuracy) and normal LLM response modes into one model. We have also improved…

open-source 3.0B ↓ 596
open-calm-7b logo
open-calm-7b
cyberagent

OpenCALM is a suite of decoder-only language models pre-trained on Japanese datasets, developed by CyberAgent, Inc.

open-source 7.0B ↓ 583
SEA-LION-v1-7B logo
SEA-LION-v1-7B
aisingapore

SEA-LION is a collection of Large Language Models (LLMs) which has been pretrained and instruct-tuned for the Southeast Asia (SEA) region. The size of the models range from 3 billion to 7 billion parameters. This is the card for the SEA-LION 7B base model.

open-source 7.0B ↓ 578