AI Agent Hub

Compare LLM Models

Select 2-5 models to compare side by side.

Select models to compare 0 / 5 selected
Select 2-5 models
Attribute
Company
Release Date 2025-10-15
Parameters (B) 8.8
Architecture Vision-Language Transformer (ViT encoder + Qwen3 dense LLM decoder with DeepStack and Interleaved-MRoPE)
Context Window 256000
Input Modalities text, image, video
Open Source Yes
License apache-2.0
Score 8.4
VRAM (GB) 18
Compute NVIDIA GPU with 16-18GB VRAM minimum for FP16/BF16 inference (~9B params plus vision encoder); 24GB recommended for comfortable runs with moderate context (e.g., RTX 4090/A6000). Q4_K_M quantization runs on 8-12GB VRAM. Apple Silicon 16GB+ unified memory or 64GB system RAM advised for local Ollama/llama.cpp deployment.
Benchmark: GSM8K 81.5
Benchmark: HumanEval 68
Benchmark: MMLU 80.7
Pricing (Input)
Pricing (Output)