Qwen: Qwen3.5-122B-A10B
About this model
Qwen3.5-122B-A10B is a native vision-language foundation model from the Qwen team, released in February 2026 as part of the Qwen3.5 family. It combines early-fusion multimodal training with a hybrid language backbone: Gated Delta Networks interleaved with gated full attention, plus a sparse mixture-of-experts stack (122 billion total parameters, about 10 billion activated per token). The design targets high-throughput inference while matching or exceeding prior Qwen3 and Qwen3-VL systems on reasoning, coding, agentic, and visual understanding tasks.
The model supports a native context length of 262,144 tokens (extendable toward roughly one million tokens with YaRN-style rope scaling) and broad multilingual coverage across 201 languages and dialects. Official benchmarks report strong knowledge scores (for example 86.7 on MMLU-Pro and 91.9 on C-Eval), competitive STEM reasoning (86.6 GPQA Diamond, 25.3 HLE with chain-of-thought), and solid agent and coding results including 72.0 on SWE-Bench Verified, 49.4 on Terminal Bench 2, and a 2100 CodeForces-style rating. On vision-language leaderboards it reaches 83.9 MMMU and 76.9 MMMU-Pro.
Weights are published under the Apache 2.0 license on Hugging Face for use with Transformers, vLLM, SGLang, and related stacks. The model runs in thinking mode by default (reasoning traces before the final answer) but can be configured for direct responses. Recommended local serving uses eight-way tensor parallelism on 80 GB-class GPUs for full 262K context; lighter quantization can reduce hardware needs at some cost to quality and context headroom.
Benchmark Scores
Technical Specs
- Parameters: 122.0B
- Architecture: Gated DeltaNet + Sparse MoE (hybrid attention)
- Context Window: 262,144 tokens
- Input Modalities: text, image
Hardware Requirements
- VRAM: 640.0 GB
- Compute: 8x NVIDIA GPU 80GB (tensor parallel, 262K context)
Pricing
| Input | Output | Currency |
|---|---|---|
| 0.26 / 1M tokens | 2.08 / 1M tokens | USD |