AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,877 models Compare
🤖
Llama-3.1-Swallow-8B-Instruct-v0.1
tokyotech-llm

Llama 3.1 Swallow is a series of large language models (8B, 70B) that were built by continual pre-training on the Meta Llama 3.1 models. Llama 3.1 Swallow enhanced the Japanese language capabilities of the original Llama 3.1 while retaining the English language capabilities. We u…

Open Source 8.0B ↓ 652
🤖
InternVL2_5-26B
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 Mini-InternVL\]](https://arxiv.org/abs/2410.16261) [\[📜 InternVL 2.5\]](https://huggingface.co/…

Multimodal 26.0B ↓ 642
🤖
Apertus-SEA-LION-v4-8B-IT
aisingapore

SEA-LION is a collection of Large Language Models (LLMs) which have been pretrained and instruct-tuned for the Southeast Asia (SEA) region.

Open Source 8.0B ↓ 636
🤖
Falcon-H1-34B-Base
tiiuae

0. TL;DR 1. Model Details 2. Training Details 3. Usage 4. Evaluation 5. Citation

Open Source 34.0B ↓ 633
🤖
MiniCPM-2B-dpo-bf16
openbmb

MiniCPM 技术报告 Technical Report OmniLMM 多模态模型 Multi-modal Model CPM-C 千亿模型试用 ~100B Model Trial

Open Source 2.0B ↓ 629
🤖
Seed-Coder-8B-Base
ByteDance-Seed

Introduction We are thrilled to introduce Seed-Coder, a powerful, transparent, and parameter-efficient family of open-source code models at the 8B scale, featuring base, instruct, and reasoning variants. Seed-Coder contributes to promote the evolution of open code models through…

Code 8.0B ↓ 627
🤖
Qwen3-Swallow-30B-A3B-RL-v0.2
tokyotech-llm

Qwen3-Swallow v0.2 is a family of large language models available in 8B , 30B-A3B , and 32B parameter sizes. Built as bilingual Japanese-English models, they were developed through Continual Pre-Training (CPT), Supervised Fine-Tuning (SFT), and Reinforcement Learning with Verifia…

Open Source 30.0B ↓ 625
🤖
HY-Embodied-0.5-X
tencent

HY-Embodied-0.5-X An Enhanced Embodied Foundation Model for Real-World Agents Tencent Robotics X × HY Vision Team

Open Source ↓ 625
🤖
UI-Mate-27B
tencent

UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations

Open Source 27.0B ↓ 620
🤖
LLaDA2.0-mini-preview
inclusionAI

LLaDA2.0-mini-preview is a diffusion language model featuring a 16BA1B Mixture-of-Experts (MoE) architecture. As an enhanced, instruction-tuned iteration of the LLaDA series, it is optimized for practical applications.

Open Source ↓ 618
🤖
Falcon-E-1B-Base
tiiuae

0. TL;DR 1. Model Details 2. Training Details 3. Usage 4. Evaluation 5. Citation

Open Source 1.0B ↓ 610
🤖
calm3-22b-chat
cyberagent

CyberAgentLM3 is a decoder-only language model pre-trained on 2.0 trillion tokens from scratch. CyberAgentLM3-Chat is a fine-tuned model specialized for dialogue use cases.

Open Source 22.0B ↓ 607