Compare LLM Models
Select 2-5 models to compare side by side.
Select models to compare
0 / 5 selected
Select 2-5 models
| Attribute | |
|---|---|
| Company | |
| Release Date | 2026-03-02 |
| Parameters (B) | 9 |
| Architecture | Hybrid Gated DeltaNet + Gated Attention (Vision-Language) |
| Context Window | 262144 |
| Input Modalities | text, image, video |
| Open Source | Yes |
| License | apache-2.0 |
| Score | 8.3 |
| VRAM (GB) | 20 |
| Compute | BF16/FP16 inference needs ~18–20 GB VRAM (e.g., RTX 4090 24 GB or L40S 48 GB); Q4_K_M quantization runs on ~6 GB VRAM (RTX 4060/3070). Supports vLLM, SGLang, Ollama, and LM Studio. |
| Benchmark: GSM8K | 89.5 |
| Benchmark: HumanEval | 82.9 |
| Benchmark: MMLU | 82.5 |
| Pricing (Input) | — |
| Pricing (Output) | — |