AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

440 models for "Base Model" Compare
Llama-3.1-Swallow-8B-Instruct-v0.5 logo
Llama-3.1-Swallow-8B-Instruct-v0.5
tokyotech-llm

Llama 3.1 Swallow is a series of large language models (8B, 70B) that were built by continual pre-training on the Meta Llama 3.1 models. Llama 3.1 Swallow enhanced the Japanese language capabilities of the original Llama 3.1 while retaining the English language capabilities. We u…

Open Source 8.0B ↓ 8.6K
Yi-1.5-6B-Chat logo
Yi-1.5-6B-Chat
01-ai

🐙 GitHub • 👾 Discord • 🐤 Twitter • 💬 WeChat 📝 Paper • 💪 Tech Blog • 🙌 FAQ • 📗 Learning Hub

Open Source 6.0B ↓ 8.3K
granite-3b-code-instruct-2k logo
granite-3b-code-instruct-2k
ibm-granite

⚠️ DEPRECATION WARNING ⚠️ ⚠️ NOT RECOMMENDED FOR USE IN NEW PROJECTS ⚠️

Code 3.0B ↓ 8.2K
InternVL3-78B logo
InternVL3-78B
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 78.4B ↓ 8.1K
LFM2.5-1.2B-JP-202606-ONNX logo
LFM2.5-1.2B-JP-202606-ONNX
LiquidAI

ONNX export of LFM2.5-1.2B-JP-202606 for cross-platform deployment via ONNX Runtime, Transformers.js, and the WebGPU stack. Same weights, same chat template — just compiled into ONNX graphs at multiple precisions, including a WebGPU-friendly INT4 + FP16 mix.

Open Source 1.17B ↓ 8.1K
Jamba-v0.1 logo
Jamba-v0.1
ai21labs

This is the base version of the Jamba model. We’ve since released a better, instruct-tuned version, Jamba-1.5-Mini. For even greater performance, check out the scaled-up Jamba-1.5-Large.

Open Source ↓ 7.6K
granite-3.0-8b-base logo
granite-3.0-8b-base
ibm-granite

Model Summary: Granite-3.0-8B-Base is a decoder-only language model to support a variety of text-to-text generation tasks. It is trained from scratch following a two-stage training strategy. In the first stage, it is trained on 10 trillion tokens sourced from diverse domains. Dur…

Open Source 8.0B ↓ 7.4K
granite-3.3-8b-base logo
granite-3.3-8b-base
ibm-granite

Granite-3.3-8B-Base is a decoder-only language model with a 128K token context window. It improves upon Granite-3.1-8B-Base by adding support for Fill-in-the-Middle (FIM) using specialized tokens, enabling the model to generate content conditioned on both prefix and suffix. This…

Open Source 8.1B ↓ 7.1K
GLM-4.5-Air-Base logo
GLM-4.5-Air-Base
zai-org

👋 Join our Discord community. 📖 Check out the GLM-4.5 technical blog , technical report , and Zhipu AI technical documentation . 📍 Use GLM-4.5 API services on Zhipu AI Open Platform . 👉 One click to GLM-4.5 .

Open Source ↓ 6.7K
InternVL3-14B logo
InternVL3-14B
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 15.1B ↓ 6.3K
Llama-3.1-Tulu-3-8B logo
Llama-3.1-Tulu-3-8B
allenai

Tülu 3 is a leading instruction following model family, offering a post-training package with fully open-source data, code, and recipes designed to serve as a comprehensive guide for modern techniques. This is one step of a bigger process to training fully open-source models, lik…

Open Source 8.0B ↓ 6.3K
OLMo-2-1124-7B-SFT logo
OLMo-2-1124-7B-SFT
allenai

Upon the initial release of OLMo-2 models, we realized the post-trained models did not share the pre-tokenization logic that the base models use. As a result, we have trained new post-trained models. The new models are available under the same names as the original models, but we…

Open Source 7.0B ↓ 6.2K