AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,308 models for "Transformer" Compare
🤖
gemma-1.1-2b-it
google

Open Source 2.0B ↓ 115.7K
🤖
LocateAnything-3B
nvidia

LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding

Multimodal 3.0B ↓ 115.1K
🤖
opt-1.3b
facebook

OPT : Open Pre-trained Transformer Language Models

Open Source 1.3B ↓ 112K
🤖
granite-vision-4.1-4b
ibm-granite

Model Summary: Granite Vision 4.1 4B is a vision-language model (VLM) that delivers frontier-level performance on structured document extraction tasks — chart extraction, table extraction, and semantic key-value pair extraction — in a compact 4B parameter footprint, providing a l…

Multimodal 4.0B ↓ 110.9K
🤖
LLaDA2.1-mini
inclusionAI

🚀 LLaDA2.1-flash is now live on ZenmuxAI ! Try it via API 🛠️ or Chat 💬: https://zenmux.ai/inclusionai/llada2.1-flash

Open Source 16.0B ↓ 108K
🤖
Cosmos-Reason1-7B
nvidia

Cosmos-Reason1: Physical AI Common Sense and Embodied Reasoning Models

Multimodal 7.0B ↓ 106.7K
🤖
biogpt
microsoft

Pre-trained language models have attracted increasing attention in the biomedical domain, inspired by their great success in the general natural language domain. Among the two main branches of pre-trained language models in the general language domain, i.e. BERT (and its variants…

Open Source 0.347B ↓ 106.4K
🤖
Llama-3.1-Swallow-8B-Instruct-v0.5
tokyotech-llm

Llama 3.1 Swallow is a series of large language models (8B, 70B) that were built by continual pre-training on the Meta Llama 3.1 models. Llama 3.1 Swallow enhanced the Japanese language capabilities of the original Llama 3.1 while retaining the English language capabilities. We u…

Open Source 8.0B ↓ 104.8K
🤖
InternVL3-8B
OpenGVLab

[\[📂 GitHub\]](https://github.com/OpenGVLab/InternVL) [\[📜 InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[📜 InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[📜 InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[📜 InternVL2.5-MPO\]](https://huggi…

Multimodal 7.94B ↓ 104.4K
🤖
SmolVLM-500M-Instruct
HuggingFaceTB

SmolVLM-500M is a tiny multimodal model, member of the SmolVLM family. It accepts arbitrary sequences of image and text inputs to produce text outputs. It's designed for efficiency. SmolVLM can answer questions about images, describe visual content, or transcribe text. Its lightw…

Multimodal 0.5B ↓ 102K
🤖
Molmo2-4B
allenai

Molmo2 is a family of open vision-language models developed by the Allen Institute for AI (Ai2) that support image, video and multi-image understanding and grounding. Molmo2 models are trained on publicly available third party datasets as referenced in our technical report and Mo…

Multimodal 4.0B ↓ 100K
🤖
ctrl
Salesforce

1. Model Details 2. Uses 3. Bias, Risks, and Limitations 4. Training 5. Evaluation 6. Environmental Impact 7. Technical Specifications 8. Citation 9. Model Card Authors 10. How To Get Started With the Model

Open Source 1.6B ↓ 99.9K