AI Agent Hub

Compare LLM Models

Select 2-5 models to compare side by side.

Select models to compare 0 / 5 selected
Select 2-5 models
Attribute
🤖 Qwen3.5-9B
Company
Release Date 2026-03-02
Parameters (B) 9
Architecture Hybrid Gated DeltaNet + Gated Attention (Vision-Language)
Context Window 262144
Input Modalities text, image, video
Open Source Yes
License apache-2.0
Score 8.3
VRAM (GB) 20
Compute BF16/FP16 inference needs ~18–20 GB VRAM (e.g., RTX 4090 24 GB or L40S 48 GB); Q4_K_M quantization runs on ~6 GB VRAM (RTX 4060/3070). Supports vLLM, SGLang, Ollama, and LM Studio.
Benchmark: GSM8K 89.5
Benchmark: HumanEval 82.9
Benchmark: MMLU 82.5
Pricing (Input)
Pricing (Output)