Compare LLM Models
Select 2-5 models to compare side by side.
Select models to compare
0 / 5 selected
Select 2-5 models
| Attribute | |
|---|---|
| Company | |
| Release Date | 2025-10-15 |
| Parameters (B) | 8.8 |
| Architecture | Vision-Language Transformer (ViT encoder + Qwen3 dense LLM decoder with DeepStack and Interleaved-MRoPE) |
| Context Window | 256000 |
| Input Modalities | text, image, video |
| Open Source | Yes |
| License | apache-2.0 |
| Score | 8.4 |
| VRAM (GB) | 18 |
| Compute | NVIDIA GPU with 16-18GB VRAM minimum for FP16/BF16 inference (~9B params plus vision encoder); 24GB recommended for comfortable runs with moderate context (e.g., RTX 4090/A6000). Q4_K_M quantization runs on 8-12GB VRAM. Apple Silicon 16GB+ unified memory or 64GB system RAM advised for local Ollama/llama.cpp deployment. |
| Benchmark: GSM8K | 81.5 |
| Benchmark: HumanEval | 68 |
| Benchmark: MMLU | 80.7 |
| Pricing (Input) | — |
| Pricing (Output) | — |