AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,316 models for "Transformer" Compare
🤖
InternVL3_5-8B-Instruct
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 8.0B ↓ 5.6K
🤖
DeepSeek-V2.5
deepseek-ai

DeepSeek-V2.5 is an upgraded version that combines DeepSeek-V2-Chat and DeepSeek-Coder-V2-Instruct. The new model integrates the general and coding abilities of the two previous versions. For model details, please visit DeepSeek-V2 page for more information.

Open Source ↓ 5.6K
🤖
DeepSeek-V3.2-Speciale
deepseek-ai

DeepSeek-V3.2: Efficient Reasoning & Agentic AI

Open Source ↓ 5.5K
🤖
OLMo-1B-0724-hf
allenai

OLMo 1B July 2024 is the latest version of the original OLMo 1B model rocking a 4.4 point increase in HellaSwag, among other evaluations improvements, from an improved version of the Dolma dataset and staged training. This version is for direct use with HuggingFace Transformers f…

Code 1.0B ↓ 5.5K
🤖
MiniMax-M1-40k
MiniMaxAI

We introduce MiniMax-M1, the world's first open-weight, large-scale hybrid-attention reasoning model. MiniMax-M1 is powered by a hybrid Mixture-of-Experts (MoE) architecture combined with a lightning attention mechanism. The model is developed based on our previous MiniMax-Text-0…

Open Source ↓ 5.5K
🤖
Apriel-Nemotron-15b-Thinker
ServiceNow-AI

1. Summary 2. Evaluation 3. Training Details 4. How to Use 5. Intended Use 6. Limitations 7. Security and Responsible Use 8. Software 9. License 10. Acknowledgements 11. Citation

Open Source 15.0B ↓ 5.5K
🤖
Flex-reddit-2x7B-1T
allenai

FlexOlmo is a new kind of LM that unlocks a new paradigm of data collaboration. With FlexOlmo, data owners can contribute to the development of open language models without giving up control of their data. There is no need to share raw data directly, and data contributors can dec…

Open Source 7.0B ↓ 5.4K
🤖
Falcon3-10B-Base
tiiuae

Falcon3 family of Open Foundation Models is a set of pretrained and instruct LLMs ranging from 1B to 10B parameters.

Open Source 10.0B ↓ 5.4K
🤖
internlm2_5-1_8b
internlm

💻Github Repo • 🤔Reporting Issues • 📜Technical Report

Open Source 8.0B ↓ 5.3K
🤖
internlm2-1_8b
internlm

💻Github Repo • 🤔Reporting Issues • 📜Technical Report

Open Source 8.0B ↓ 5.3K
🤖
GLM-4.6V-FP8
zai-org

This model is part of the GLM-V family of models, introduced in the paper GLM-4.1V-Thinking and GLM-4.5V: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

Open Source ↓ 5.3K
🤖
InternVL2_5-1B
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 Mini-InternVL\]](https://arxiv.org/abs/2410.16261) [\[📜 InternVL 2.5\]](https://huggingface.co/…

Multimodal 1.0B ↓ 5.2K