Qwen: Qwen3.8 Max Prime
About this model
Qwen3.8 Max Prime is a hosted serving SKU from Alibaba's Qwen team, released in September 2026. It runs the same underlying weights and capabilities as Qwen3.8 Max—the 2.4-trillion-parameter mixture-of-experts flagship with roughly 95 billion active parameters per forward pass—while targeting higher throughput and capacity on Alibaba's inference stack. The endpoint is priced at roughly double the standard Max tier and is aimed at latency-sensitive coding agents, multi-agent orchestration, and high-concurrency production workloads rather than a separate training run or upgraded model family.
The model supports a one-million-token context window, reasoning enabled by default with configurable reasoning effort, tool calling, and structured outputs. Inputs include text, images, and video; outputs are text. Benchmarks published for Qwen3.8 Max (inherited by Prime) emphasize long-horizon software engineering and agentic work—strong Terminal-Bench 2.1, SWE-Bench Pro, NL2Repo-Bench, and multimodal MMMU-Pro scores—alongside competitive GPQA-Diamond, MMLU-Pro, and instruction-following results. Independent leaderboards also report solid LiveCodeBench and SWE-Bench Verified performance for the Max weight snapshot.
Max Prime does not ship open weights; self-hosting is available via the separate Qwen3.8-2.4T-A95B release under the Qwen3.8-Max License, while Prime remains API-only through providers such as OpenRouter and Qwen Cloud. Teams that are not throughput-bound should prefer the standard qwen3.8-max or pinned qwen3.8-max-0902 endpoints at lower cost for identical model behavior.
Benchmark Scores
Technical Specs
- Parameters: 2400.0B
- Architecture: Mixture-of-Experts Transformer
- Context Window: 1,000,000 tokens
- Input Modalities: text, image, video
Hardware Requirements
- Compute: API only
Pricing
| Input | Output | Currency |
|---|---|---|
| 4.00 / 1M tokens | 12.00 / 1M tokens | USD |