AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,115 models for "Chat" Compare
🤖
xLAM-1b-fc-r
Salesforce

[Homepage] [APIGen Paper] [ActionStudio Paper] [Discord] [Dataset] [Github]

Open Source 1.0B ↓ 4K
🤖
AprielGuard
ServiceNow-AI

1. Summary 2. Taxonomy 2. Evaluation 3. Training Details 4. How to Use 5. Intended Use 6. Limitations 7. License 8. Citation

Open Source ↓ 4K
🤖
Ring-mini-linear-2.0
inclusionAI

📖 Technical Report &nbsp&nbsp &nbsp&nbsp 🤗 Hugging Face &nbsp&nbsp &nbsp&nbsp🤖 ModelScope

Open Source ↓ 4K
🤖
Qwen-SEA-LION-v4.5-27B-IT
aisingapore

SEA-LION is a collection of Large Language Models (LLMs) which have been pretrained and instruct-tuned for the Southeast Asia (SEA) region.

Open Source 27.0B ↓ 3.9K
🤖
Hermes-4-14B
NousResearch

\ud83d\udcda Paper (Hugging Face) \ud83d\udcda Paper (arXiv) \ud83c\udf10 Project Page \ud83d\udcbb GitHub Repository

Open Source 14.0B ↓ 3.9K
🤖
InternVL3-9B
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 9.0B ↓ 3.9K
🤖
LFM2.5-1.2B-JP
LiquidAI

LFM2.5-1.2B-JP is a chat model specifically optimized for Japanese. While LFM2 already supported Japanese as one of eight languages, LFM2.5-JP pushes state-of-the-art on Japanese knowledge and instruction-following at its scale. This model is ideal for developers building Japanes…

Open Source 1.2B ↓ 3.8K
🤖
Hermes-2-Pro-Llama-3-8B
NousResearch

Hermes 2 Pro is an upgraded, retrained version of Nous Hermes 2, consisting of an updated and cleaned version of the OpenHermes 2.5 Dataset, as well as a newly introduced Function Calling and JSON Mode dataset developed in-house.

Open Source 8.0B ↓ 3.8K
🤖
LLaDA2.0-flash
inclusionAI

LLaDA2.0-flash is a diffusion language model featuring a 100BA6B Mixture-of-Experts (MoE) architecture. As an enhanced, instruction-tuned iteration of the LLaDA2.0 series, it is optimized for practical applications.

Open Source ↓ 3.7K
🤖
Qwen3-Swallow-8B-CPT-v0.2
tokyotech-llm

Qwen3-Swallow v0.2 is a family of large language models available in 8B , 30B-A3B , and 32B parameter sizes. Built as bilingual Japanese-English models, they were developed through Continual Pre-Training (CPT), Supervised Fine-Tuning (SFT), and Reinforcement Learning with Verifia…

Open Source 8.0B ↓ 3.6K
🤖
OLMo-2-1124-7B-DPO
allenai

Upon the initial release of OLMo-2 models, we realized the post-trained models did not share the pre-tokenization logic that the base models use. As a result, we have trained new post-trained models. The new models are available under the same names as the original models, but we…

Open Source 7.0B ↓ 3.5K
🤖
falcon-11B
tiiuae

Falcon2-11B is an 11B parameters causal decoder-only model built by TII and trained on over 5,000B tokens of RefinedWeb enhanced with curated corpora. The model is made available under the TII Falcon License 2.0, the permissive Apache 2.0-based software license which includes an…

Open Source 11.0B ↓ 3.5K