AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,315 models for "Transformer" Compare
🤖
xLAM-1b-fc-r
Salesforce

[Homepage] [APIGen Paper] [ActionStudio Paper] [Discord] [Dataset] [Github]

Open Source 1.0B ↓ 4K
🤖
codegen-350M-multi
Salesforce

CodeGen is a family of autoregressive language models for program synthesis from the paper: A Conversational Paradigm for Program Synthesis by Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, Caiming Xiong. The models are originally releas…

Code ↓ 4K
🤖
AprielGuard
ServiceNow-AI

1. Summary 2. Taxonomy 2. Evaluation 3. Training Details 4. How to Use 5. Intended Use 6. Limitations 7. License 8. Citation

Open Source ↓ 4K
🤖
Ring-mini-linear-2.0
inclusionAI

📖 Technical Report &nbsp&nbsp &nbsp&nbsp 🤗 Hugging Face &nbsp&nbsp &nbsp&nbsp🤖 ModelScope

Open Source ↓ 4K
🤖
Qwen-SEA-LION-v4.5-27B-IT
aisingapore

SEA-LION is a collection of Large Language Models (LLMs) which have been pretrained and instruct-tuned for the Southeast Asia (SEA) region.

Open Source 27.0B ↓ 3.9K
🤖
Hermes-4-14B
NousResearch

\ud83d\udcda Paper (Hugging Face) \ud83d\udcda Paper (arXiv) \ud83c\udf10 Project Page \ud83d\udcbb GitHub Repository

Open Source 14.0B ↓ 3.9K
🤖
InternVL3-9B
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 9.0B ↓ 3.9K
🤖
LFM2.5-1.2B-JP
LiquidAI

LFM2.5-1.2B-JP is a chat model specifically optimized for Japanese. While LFM2 already supported Japanese as one of eight languages, LFM2.5-JP pushes state-of-the-art on Japanese knowledge and instruction-following at its scale. This model is ideal for developers building Japanes…

Open Source 1.2B ↓ 3.8K
🤖
Hermes-2-Pro-Llama-3-8B
NousResearch

Hermes 2 Pro is an upgraded, retrained version of Nous Hermes 2, consisting of an updated and cleaned version of the OpenHermes 2.5 Dataset, as well as a newly introduced Function Calling and JSON Mode dataset developed in-house.

Open Source 8.0B ↓ 3.8K
🤖
Falcon3-3B-Base
tiiuae

Falcon3 family of Open Foundation Models is a set of pretrained and instruct LLMs ranging from 1B to 10B parameters.

Open Source 3.0B ↓ 3.7K
🤖
LLaDA2.0-flash
inclusionAI

LLaDA2.0-flash is a diffusion language model featuring a 100BA6B Mixture-of-Experts (MoE) architecture. As an enhanced, instruction-tuned iteration of the LLaDA2.0 series, it is optimized for practical applications.

Open Source ↓ 3.7K
🤖
granite-3.1-1b-a400m-base
ibm-granite

Model Summary: Granite-3.1-1B-A400M-Base extends the context length of Granite-3.0-1B-A400M-Base from 4K to 128K using a progressive training strategy by increasing the supported context length in increments while adjusting RoPE theta until the model has successfully adapted to d…

Open Source 1.0B ↓ 3.7K