AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

440 models for "Base Model" Compare
AHN-GDN-for-Qwen-2.5-Instruct-14B logo
AHN-GDN-for-Qwen-2.5-Instruct-14B
ByteDance-Seed

AHN: Artificial Hippocampus Networks for Efficient Long-Context Modeling

Open Source 14.0B ↓ 134
Klear-46B-A2.5B-Instruct logo
Klear-46B-A2.5B-Instruct
Kwai-Klear

🤗 Hugging Face 💻 Github Repository 📑 Technique Report 💬 Issues & Discussions

Open Source 46.0B ↓ 133
bloom-3b-intermediate logo
bloom-3b-intermediate
bigscience

WARNING: The checkpoints on this repo are not fully trained model. Evaluations of intermediary checkpoints and the final model will be added when conducted (see below).

Open Source 3.0B ↓ 127
AHN-DN-for-Qwen-2.5-Instruct-3B logo
AHN-DN-for-Qwen-2.5-Instruct-3B
ByteDance-Seed

AHN: Artificial Hippocampus Networks for Efficient Long-Context Modeling

Open Source 3.0B ↓ 126
Klear-46B-A2.5B-Base logo
Klear-46B-A2.5B-Base
Kwai-Klear

🤗 Hugging Face 💻 Github Repository 📑 Technique Report 💬 Issues & Discussions

Open Source 46.0B ↓ 123
Swallow-13b-NVE-hf logo
Swallow-13b-NVE-hf
tokyotech-llm

Our Swallow model has undergone continual pre-training from the Llama 2 family, primarily with the addition of Japanese language data. The tuned versions use supervised fine-tuning (SFT). Links to other models can be found in the index.

Open Source 13.0B ↓ 120
bloom-1b7-intermediate logo
bloom-1b7-intermediate
bigscience

WARNING: The checkpoints on this repo are not fully trained model. Evaluations of intermediary checkpoints and the final model will be added when conducted (see below).

Open Source ↓ 120
AHN-DN-for-Qwen-2.5-Instruct-7B logo
AHN-DN-for-Qwen-2.5-Instruct-7B
ByteDance-Seed

AHN: Artificial Hippocampus Networks for Efficient Long-Context Modeling

Open Source 7.0B ↓ 117
Baichuan-M3-235B-FP8 logo
Baichuan-M3-235B-FP8
baichuan-inc

From Inquiry to Decision: Building Trustworthy Medical AI

Open Source 235.0B ↓ 109
SEA-LION-v1-7B-IT-Research logo
SEA-LION-v1-7B-IT-Research
aisingapore

SEA-LION is a collection of Large Language Models (LLMs) which has been pretrained and instruct-tuned for the Southeast Asia (SEA) region. The size of the models range from 3 billion to 7 billion parameters. This is the card for the SEA-LION 7B Instruct (Non-Commercial) model.

Open Source 7.0B ↓ 107
Qwen-SEA-LION-v4-32B-IT-OV-8BIT logo
Qwen-SEA-LION-v4-32B-IT-OV-8BIT
aisingapore

- Model creator: AI Singapore - Original model: Qwen-SEA-LION-v4-32B-IT

Open Source 32.0B ↓ 104
mistral-coreml logo
mistral-coreml
apple

[!IMPORTANT] ❗ This repo requires the use of the macOS Sequoia (15) Developer Beta to utilize the latest and greatest CoreML has to offer! Sign up for the Apple Beta Software Program here to get access. Check out the companion blog post to learn more about what's new in iOS 18 &…

Open Source ↓ 93