Compare LLM Models
Select 2-5 models to compare side by side.
Select models to compare
0 / 5 selected
Select 2-5 models
| Attribute | |
|---|---|
| Company | |
| Release Date | 2025-01-26 |
| Parameters (B) | 7 |
| Architecture | Vision-Language Transformer (ViT + Qwen2.5 LLM with MRoPE) |
| Context Window | 32768 |
| Input Modalities | text, image |
| Open Source | Yes |
| License | apache-2.0 |
| Score | 8.3 |
| VRAM (GB) | 16 |
| Compute | BF16 inference needs ~13 GB VRAM minimum (official); ~16 GB recommended with KV cache at moderate context. RTX 4090 24GB or A100 40GB ideal for production. INT4/AWQ quantization reduces requirement to ~4.5 GB. Flash Attention 2 strongly recommended for multi-image/video workloads. |
| Benchmark: GSM8K | 91.6 |
| Benchmark: HumanEval | 84.8 |
| Benchmark: MMLU | 71.1 |
| Pricing (Input) | — |
| Pricing (Output) | — |