AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,304 models for "Transformer" Compare
🤖
Meta-Llama-3-8B
meta-llama

Open Source 8.0B ↓ 627.2K
🤖
Phi-3-mini-4k-instruct
microsoft

🎉 Phi-3.5 : [[mini-instruct]](https://huggingface.co/microsoft/Phi-3.5-mini-instruct); [[MoE-instruct]](https://huggingface.co/microsoft/Phi-3.5-MoE-instruct) ; [[vision-instruct]](https://huggingface.co/microsoft/Phi-3.5-vision-instruct)

Open Source ↓ 612.3K
🤖
InternVL2-2B
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 Mini-InternVL\]](https://arxiv.org/abs/2410.16261) [\[📜 InternVL 2.5\]](https://huggingface.co/…

Multimodal 2.2B ↓ 608.1K
🤖
InternVL2-1B
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 Mini-InternVL\]](https://arxiv.org/abs/2410.16261) [\[📜 InternVL 2.5\]](https://huggingface.co/…

Multimodal 0.9B ↓ 592.9K
🤖
SmolLM3-3B
HuggingFaceTB

1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. License

Open Source 3.0B ↓ 585.8K
🤖
DeepSeek-R1-Distill-Qwen-32B
deepseek-ai

We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With R…

Reasoning 32.0B ↓ 577.8K
🤖
SmolVLM-256M-Instruct
HuggingFaceTB

SmolVLM-256M is the smallest multimodal model in the world. It accepts arbitrary sequences of image and text inputs to produce text outputs. It's designed for efficiency. SmolVLM can answer questions about images, describe visual content, or transcribe text. Its lightweight archi…

Multimodal 0.256B ↓ 577K
🤖
falcon-7b
tiiuae

Falcon-7B is a 7B parameters causal decoder-only model built by TII and trained on 1,500B tokens of RefinedWeb enhanced with curated corpora. It is made available under the Apache 2.0 license.

Open Source 7.0B ↓ 563.1K
🤖
Kimi-K2.5
moonshotai

📰   Tech Blog         📄   Paper

Open Source ↓ 554.9K
🤖
japanese-gpt-neox-small
rinna

This repository provides a small-sized Japanese GPT-NeoX model. The model was trained using code based on EleutherAI/gpt-neox.

Open Source 0.11B ↓ 548.1K
🤖
Cosmos-Reason2-8B
nvidia

Multimodal 8.0B ↓ 527K
🤖
DeepSeek-R1-Distill-Qwen-1.5B
deepseek-ai

We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With R…

Reasoning 1.5B ↓ 488.7K