AI Agent Hub

Compare LLM Models

Select 2-5 models to compare side by side.

Select models to compare 0 / 5 selected
Select 2-5 models
Attribute
Company
Release Date 2025-01-26
Parameters (B) 7
Architecture Vision-Language Transformer (ViT + Qwen2.5 LLM with MRoPE)
Context Window 32768
Input Modalities text, image
Open Source Yes
License apache-2.0
Score 8.3
VRAM (GB) 16
Compute BF16 inference needs ~13 GB VRAM minimum (official); ~16 GB recommended with KV cache at moderate context. RTX 4090 24GB or A100 40GB ideal for production. INT4/AWQ quantization reduces requirement to ~4.5 GB. Flash Attention 2 strongly recommended for multi-image/video workloads.
Benchmark: GSM8K 91.6
Benchmark: HumanEval 84.8
Benchmark: MMLU 71.1
Pricing (Input)
Pricing (Output)