AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,308 models for "Transformer" Compare
🤖
granite-guardian-3.3-8b
ibm-granite

Model Summary: Granite Guardian 3.3 8b is a specialized Granite 3.3 8B model designed to judge if the input prompts and the output responses of an LLM based system meet specified criteria. The model comes pre-baked with certain criteria including but not limited to: jailbreak att…

Open Source 8.0B ↓ 217.6K
🤖
olmOCR-2-7B-1025
allenai

Full BF16 version of olmOCR-2-7B-1025-FP8. We recommend using the FP8 version for all practical purposes except further fine tuning.

Multimodal 7.0B ↓ 213.8K
🤖
granite-docling-258M
ibm-granite

granite-docling-258m Granite Docling is a multimodal Image-Text-to-Text model engineered for efficient document conversion. It preserves the core features of Docling while maintaining seamless integration with DoclingDocuments to ensure full compatibility.

Multimodal 0.258B ↓ 212.6K
🤖
gemma-3n-E2B-it
google

Multimodal 2.0B ↓ 204.9K
🤖
InternVL3-1B
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 0.94B ↓ 200.8K
🤖
Llama-Guard-3-8B
meta-llama

Open Source 8.0B ↓ 196K
🤖
opt-350m
facebook

OPT : Open Pre-trained Transformer Language Models

Open Source ↓ 191.2K
🤖
paligemma-3b-pt-224
google

Multimodal 3.0B ↓ 190.8K
🤖
SmolLM2-1.7B-Instruct
HuggingFaceTB

1. Model Summary 2. Evaluation 3. Examples 4. Limitations 5. Training 6. License 7. Citation

Open Source 1.7B ↓ 185.2K
🤖
Mistral-7B-Instruct-v0.1
mistralai

py from mistral common.tokens.tokenizers.mistral import MistralTokenizer from mistral common.protocol.instruct.messages import UserMessage from mistral common.protocol.instruct.request import ChatCompletionRequest

Open Source 7.0B ↓ 178.1K
🤖
InternVL2-26B
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 Mini-InternVL\]](https://arxiv.org/abs/2410.16261) [\[📜 InternVL 2.5\]](https://huggingface.co/…

Multimodal 25.5B ↓ 177.1K
🤖
Llama-3_3-Nemotron-Super-49B-v1
nvidia

Llama-3.3-Nemotron-Super-49B-v1 is a large language model (LLM) which is a derivative of Meta Llama-3.3-70B-Instruct (AKA the reference model ). It is a reasoning model that is post trained for reasoning, human chat preferences, and tasks, such as RAG and tool calling. The model…

Open Source 49.0B ↓ 176.5K