AI Agent Hub

LLM Models

Browse the world's large language models. Compare parameters, benchmarks, VRAM and more.

1,304 models for "Transformer" Compare
🤖
Llama-3.1-Nemotron-Nano-VL-8B-V1
nvidia

Llama Nemotron Nano VL is a leading document intelligence vision language model (VLMs) that enables the ability to query and summarize images from the physical or virtual world. Llama Nemotron Nano VL is deployable in the data center, cloud and at the edge, including Jetson Orin…

Multimodal 9.0B ↓ 301.4K
🤖
Phi-3.5-mini-instruct
microsoft

🎉 Phi-4 : [multimodal-instruct onnx]; [mini-instruct onnx]

Open Source ↓ 300.7K
🤖
Mage-VL
microsoft

Mage-VL An Efficient Codec-Native Streaming Multimodal Foundation Model

Multimodal 4.0B ↓ 293K
🤖
gemma-2-2b
google

Open Source 2.0B ↓ 292.9K
🤖
Meta-Llama-3.1-8B-Instruct
NousResearch

The Meta Llama 3.1 collection of multilingual large language models (LLMs) is a collection of pretrained and instruction tuned generative models in 8B, 70B and 405B sizes (text in/text out). The Llama 3.1 instruction tuned text only models (8B, 70B, 405B) are optimized for multil…

Open Source 8.0B ↓ 292.7K
🤖
Phi-3-mini-128k-instruct
microsoft

🎉 Phi-4 : [multimodal-instruct onnx]; [mini-instruct onnx]

Open Source ↓ 286.9K
🤖
medgemma-27b-it
google

Multimodal 27.0B ↓ 277.1K
🤖
Phi-tiny-MoE-instruct
microsoft

Phi-tiny-MoE is a lightweight Mixture of Experts (MoE) model with 3.8B total parameters and 1.1B activated parameters. It is compressed and distilled from the base model shared by Phi-3.5-MoE and GRIN-MoE using the SlimMoE approach, then post-trained via supervised fine-tuning an…

Open Source ↓ 272.7K
🤖
deepseek-vl2-tiny
deepseek-ai

Introducing DeepSeek-VL2, an advanced series of large Mixture-of-Experts (MoE) Vision-Language Models that significantly improves upon its predecessor, DeepSeek-VL. DeepSeek-VL2 demonstrates superior capabilities across various tasks, including but not limited to visual question…

Multimodal 3.37B ↓ 265.9K
🤖
OLMoE-1B-7B-0125-Instruct
allenai

OLMoE-1B-7B-0125-Instruct January 2025 is post-trained variant of the OLMoE-1B-7B January 2025 model, which has undergone supervised finetuning on an OLMo-specific variant of the Tülu 3 dataset and further DPO training on this dataset, and finally RLVR training using this data. T…

Open Source 1.0B ↓ 264.8K
🤖
medgemma-1.5-4b-it
google

Multimodal 4.0B ↓ 263.8K
🤖
olmOCR-2-7B-1025-FP8
allenai

Quantized to FP8 Version of olmOCR-2-7B-1025, using llmcompressor.

Multimodal 7.0B ↓ 260K