AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,308 models for "Transformer" Compare
🤖
stablelm-base-alpha-7b
stabilityai

📢 DISCLAIMER : The StableLM-Base-Alpha models have been superseded. Find the latest versions in the Stable LM Collection here.

Open Source 7.0B ↓ 1.1K
🤖
mt0-base
bigscience

1. Model Summary 2. Use 3. Limitations 4. Training 5. Evaluation 7. Citation

Open Source ↓ 1.1K
🤖
internlm2_5-20b-chat
internlm

💻Github Repo • 🤔Reporting Issues • 📜Technical Report

Open Source 20.0B ↓ 1.1K
🤖
RedPajama-INCITE-Base-3B-v1
togethercomputer

RedPajama-INCITE-Base-3B-v1 was developed by Together and leaders from the open-source AI community including Ontocord.ai, ETH DS3Lab, AAI CERC, Université de Montréal, MILA - Québec AI Institute, Stanford Center for Research on Foundation Models (CRFM), Stanford Hazy Research re…

Open Source 3.0B ↓ 1.1K
🤖
bloomz-1b1
bigscience

1. Model Summary 2. Use 3. Limitations 4. Training 5. Evaluation 7. Citation

Open Source ↓ 1K
🤖
Falcon-H1-7B-Base
tiiuae

0. TL;DR 1. Model Details 2. Training Details 3. Usage 4. Evaluation 5. Citation

Open Source 7.0B ↓ 1K
🤖
InternVL3_5-241B-A28B-Instruct
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 241.0B ↓ 1K
🤖
GLM-4-32B-Base-0414
zai-org

The GLM family welcomes new members, the GLM-4-32B-0414 series models, featuring 32 billion parameters. Its performance is comparable to OpenAI’s GPT series and DeepSeek’s V3/R1 series. It also supports very user-friendly local deployment features. GLM-4-32B-Base-0414 was pre-tra…

Open Source 32.0B ↓ 1K
🤖
japanese-gpt-neox-3.6b-instruction-sft-v2
rinna

japanese-gpt-neox-3.6b-instruction-sft-v2

Open Source 3.6B ↓ 1K
🤖
Nous-Hermes-2-Mistral-7B-DPO
NousResearch

Nous Hermes 2 on Mistral 7B DPO is the new flagship 7B Hermes! This model was DPO'd from Teknium/OpenHermes-2.5-Mistral-7B and has improved across the board on all benchmarks tested - AGIEval, BigBench Reasoning, GPT4All, and TruthfulQA.

Open Source 7.0B ↓ 1K
🤖
RedPajama-INCITE-Instruct-3B-v1
togethercomputer

RedPajama-INCITE-Instruct-3B-v1 was developed by Together and leaders from the open-source AI community including Ontocord.ai, ETH DS3Lab, AAI CERC, Université de Montréal, MILA - Québec AI Institute, Stanford Center for Research on Foundation Models (CRFM), Stanford Hazy Researc…

Open Source 3.0B ↓ 988
🤖
codegen-2B-multi
Salesforce

CodeGen is a family of autoregressive language models for program synthesis from the paper: A Conversational Paradigm for Program Synthesis by Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, Caiming Xiong. The models are originally releas…

Code 2.0B ↓ 985