AI Agent Hub
Back to models
Qwen: Qwen3.8 Omni Flash logo

Qwen: Qwen3.8 Omni Flash

Closed Source qwen Released 2026-09-20
-- 125.0B params 1M context Proprietary

About this model

Qwen3.8-Omni-Flash is Alibaba Qwen's native omnimodal agent model for real-world productivity. It unifies text, image, audio, and video in a single Thinker-Talker sparse MoE stack built on the Qwen3.8-Next language backbone, with upgraded AuT and Spatial AuT encoders for general and spatial audio. The model supports up to one million tokens of context for long-form audiovisual reasoning, meeting understanding, and multi-step agent planning while keeping text performance in line with Qwen3.8-Flash-class models.

Training combines native multimodal co-training, specialist distillation, and reinforcement learning on long-horizon agent data so reasoning and tool-use skills transfer from text to vision and audio-visual tasks. Evaluations in the Qwen3.8-Omni technical report highlight strong coding and agent scores (for example SWE-Bench-Pro 63.3, LiveCodeBench v6 92.6, GPQA Diamond 91.0, HLE 36.5) alongside large gains over Qwen3.5-Omni-Plus on audio, video, and multimodal agent benchmarks.

The model is offered primarily through DashScope and Qwen Cloud APIs rather than as a fully open weight release. Alibaba pairs it with Qwen-MM-Plugins and Qwen-Live-Harness so developers can plug audiovisual perception, selective retrieval, and realtime interaction into existing agent frameworks for workflows such as video editing, localization, film commentary, and omnimodal research.

Benchmark Scores

HLE
36.5
GPQA
91.0
DeepSWE
57.8
GPQA-Diamond
91.0
NL2Repo-Bench
48.9

Technical Specs

  • Parameters: 125.0B
  • Architecture: Sparse MoE Transformer (Thinker-Talker)
  • Context Window: 1,000,000 tokens
  • Input Modalities: text, image, audio

Hardware Requirements

  • Compute: API only

Pricing

Input Output Currency
0.15 / 1M tokens 0.47 / 1M tokens USD