AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,311 models for "Transformer" Compare
🤖
Kimi-VL-A3B-Thinking
moonshotai

[!Warning] This model has a new version: Kimi-VL-A3B-Thinking-2506. Please consider using the new 2506 version for better abilties on general visual understanding, reasoning, video and agent scenarios.

Multimodal 16.0B ↓ 45.9K
🤖
granite-3.1-8b-instruct
ibm-granite

Model Summary: Granite-3.1-8B-Instruct is a 8B parameter long-context instruct model finetuned from Granite-3.1-8B-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets tailored for solving long context pr…

Open Source 8.0B ↓ 45.3K
🤖
open-calm-3b
cyberagent

OpenCALM is a suite of decoder-only language models pre-trained on Japanese datasets, developed by CyberAgent, Inc.

Open Source 2.7B ↓ 45.2K
🤖
Llama-3.1-70B
meta-llama

Open Source 70.0B ↓ 45.2K
🤖
MiMo-7B-RL
XiaomiMiMo

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ Unlocking the Reasoning Potential of Language Model From Pretraining to Posttraining ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

Open Source 7.0B ↓ 44.8K
🤖
MiniCPM4-0.5B
openbmb

GitHub Repo Technical Report 👋 Join us on Discord and WeChat

Open Source 0.5B ↓ 43.6K
🤖
Molmo2-O-7B
allenai

Molmo2 is a family of open vision-language models developed by the Allen Institute for AI (Ai2) that support image, video and multi-image understanding and grounding. Molmo2 models are trained on publicly available third party datasets as referenced in our technical report and Mo…

Multimodal 7.0B ↓ 43.5K
🤖
Emu3-Chat-hf
BAAI

Emu3: Next-Token Prediction is All You Need

Multimodal 8.0B ↓ 41.7K
🤖
InternVL2-8B
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 Mini-InternVL\]](https://arxiv.org/abs/2410.16261) [\[📜 InternVL 2.5\]](https://huggingface.co/…

Multimodal 8.1B ↓ 41.6K
🤖
starcoder2-7b
bigcode

1. Model Summary 2. Use 3. Limitations 4. Training 5. License 6. Citation

Code 7.0B ↓ 41.3K
🤖
gpt_bigcode-santacoder
bigcode

Play with the model on the SantaCoder Space Demo.

Code 1.1B ↓ 41.3K
🤖
Meta-Llama-3.1-8B
NousResearch

The Meta Llama 3.1 collection of multilingual large language models (LLMs) is a collection of pretrained and instruction tuned generative models in 8B, 70B and 405B sizes (text in/text out). The Llama 3.1 instruction tuned text only models (8B, 70B, 405B) are optimized for multil…

Open Source 8.0B ↓ 41.1K