Foreword¶
On July 19, 2026, during the Shanghai WAIC (World Artificial Intelligence Conference), Alibaba’s Qwen team announced Qwen3.8-Max-Preview on X: with a total parameter count of 2.4 trillion, it supports multimodal inputs, claims to be “second only to Claude Fable 5” in capability, and promises to release the weights “soon”. The preview endpoint is already available via Token Plan, Qoder, and QoderWork.
Unlike Kimi K3, which released a benchmark table, pricing, and a July 27 weight release date when it launched on July 16, Qwen 3.8’s launch came with almost just one tweet: no benchmarks, no model card, no formal token-based API pricing. The developer community quickly shifted its focus from “how strong is the model” to a more concrete question: Does “soon” count as a date?
Based on Qwen’s official announcements, Qwen Cloud documentation, and cross-verification from multiple media outlets, this article sorts out the currently confirmed facts, the still unresolved metrics, and how ordinary developers should rationally trial this preview version.
July 19: A “Preview” at WAIC, Not an Official Launch¶
According to Qwen’s official X post and reports from media outlets including Bloomberg and Quartz, Qwen3.8-Max-Preview was publicly announced on July 19, 2026 (Sunday), with the launch coinciding with the WAIC 2026 time window. It is important to emphasize: what is currently live is the preview endpoint, not the final production version, let alone a downloadable open-weight checkpoint.
Qwen Cloud’s documentation states that the model will continue to iterate during the preview period, and may be taken offline or replaced by the official version later. Therefore, results obtained today may not be reproducible next week—this is a hard constraint when evaluating frontier models.
The official product identifier is:
qwen3.8-max-preview
Be sure to use the full model ID when connecting to avoid confusion with the older Qwen3-8B (an 8-billion-parameter small model). Qwen3.8 is a new generational naming, and 2.4T is its flagship preview specification.
2.4T MoE and Multimodality: Confirmed and Undisclosed¶
Verified Information¶
| Item | Current Status |
|---|---|
| Total Parameters | 2.4 trillion (official Qwen口径) |
| Architecture | Sparse Mixture-of-Experts (MoE), consistently cited by media and third-party reviews |
| Multimodality | Supports text, image, video, and document inputs (consistent with official reports and coverage from GIGAZINE, TechTimes, etc.) |
| Preview Access | Token Plan, Qoder, QoderWork |
| Weight Release | Promised “soon”, no date, no license, no Hugging Face repository |
Qwen states that this is the first flagship preview in the Qwen family with total parameters exceeding 1T and multimodal capabilities, which indeed marks a leap in scale.
Still Unreleased Key Metrics¶
As of July 29, 2026, Qwen has not disclosed these in an official model card or technical report:
1. Active parameters per token — determines inference costs and actual MoE computing power consumption;
2. Complete benchmark tables and evaluation methodology;
3. Final open-weight license (whether Apache 2.0 or Qwen Research License, etc.);
4. Standard API pricing per million tokens (the preview period uses Credits subscription, see below).
Third-party documents (such as the Qwen Cloud Codex integration guide) list some integration layer parameters: context window of approximately 983,616 tokens, maximum output of 131,072 tokens, inference mode defaults to xhigh and cannot be turned off. These are integration specifications in the access documentation, cannot replace the official model card, but are more reliable than the second-hand claim of “about 1M context” circulating online.
“Second Only to Fable 5”: A Positioning Statement, No Tables Included¶
The original English wording in Qwen’s post was: compatible to leading frontier AI models, second only to Fable 5. Multiple media outlets (Quartz, TechTimes, GIGAZINE) have pointed out: the announcement did not include any benchmark scores, evaluation configurations, or reproduction scripts.
Compare this to the previous generation Qwen3.7-Max (May 2026): at that time, Alibaba released relatively complete technical materials, such as a score of 56.6 on the Artificial Analysis Intelligence Index. While Qwen3.7 itself did not release weights, at least there was data available for reference.
Qwen3.8 even skipped that step — as of the week of its launch, there was no independent ranking entry for Qwen3.8 on third-party leaderboards such as LMArena; Claude Fable 5 still tops the list, and the previous generation Qwen3.7-Max ranks significantly lower. Therefore:
The “world’s second” position is currently the manufacturer’s self-positioning, not a reproducible evaluation conclusion.
Independent tests have emerged sporadically (such as matched tests on tasks like StackPerf by some bloggers), but the sample size is small, and the preview endpoint itself is changing, so they cannot be used as a replacement for official benchmarks.
Is “Soon” a Date: Where the Open-Source Suspense Comes From¶
This is the most controversial point in the Chinese developer community during this launch.
Qwen’s Historical “Open” and “Closed” Strategies¶
Verifying Qwen’s flagship strategy for recent generations, the timeline is roughly as follows:
| Generation | Time | Weight Strategy | Release Material Completeness |
|---|---|---|---|
| Qwen3 and earlier flagships | Pre-2026 | Most provided downloadable weights, mostly Apache 2.0 | Relatively complete |
| Qwen3.6-Max-Preview | 2026-04-20 | First closed flagship, only API access | Had benchmarks and pricing |
| Qwen3.7-Max | 2026-05 | Closed, no weight release promised | Had technical documentation and metrics |
| Qwen3.8-Max-Preview | 2026-07-19 | Preview API + promise of open-weight “soon” | No benchmarks, no date |
Qwen3.6 and 3.7 never promised open-source for their Max flagships; Qwen3.8 is the first Max-level model in the closed-preview era to revive the open-weight narrative again. The community is therefore both excited and wary: The words are back, but the calendar is not.
Why “Soon” Sparks Debate¶
AI commentator Julien Simon summarized the differences between the launches of Kimi K3 and Qwen 3.8 as “dated gap vs undated gap”:
- K3: Had benchmarks, a pricing of $15/M output tokens, and a weight release promise of July 27 — failure to deliver by the deadline is a falsifiable event;
- Qwen 3.8: Also said open-weight soon, but no deadline — “soon” can only drift, and cannot be “falsified”.
For developers, this means:
1. Do not treat “soon” as an implied release date for scheduling;
2. Do not call it an open-source model until an official repository with a license file appears on Hugging Face;
3. Even if weights are released in the future, the 2.4 trillion total parameters will require approximately 1.2TB of VRAM under 4-bit quantization (estimated by TechTimes), and multi-node or more aggressive quantization will be necessary — without disclosing active parameters, self-hosting costs still cannot be accurately calculated.
Context: “Same-Week Narrative” Three Days After K3’s Launch¶
The timeline is worth noting separately: Kimi K3 launched on July 16, and Qwen 3.8 followed up on July 19. K3 was called by multiple media outlets one of the highest-priced frontier models from Chinese manufacturers to date (output at approximately $15 per million tokens), and gave a July 27 open-weight release date.
Bloomberg reported that K3’s launch disturbed global tech stock sentiment; after Qwen 3.8’s announcement, Alibaba’s stock rose approximately 5.4% in the next trading day (Quartz data). Additionally, public documents show that Alibaba acquired approximately 36% of Moonshot (parent company of Kimi) for approximately $800 million (2024 fiscal year 20-F report, which may have been diluted in subsequent rounds). The competitive and cooperative relationship between the two companies provides a capital market perspective for the “same-week dual flagship narrative”, but this does not change the fundamental principle that technical evaluation should be based on reproducible evidence.
How to Trial Qwen3.8-Max-Preview¶
If you want to experience the preview version under controlled conditions, you can follow these steps.
1. Token Plan Subscription (Main Path for Individual Developers)¶
The current promotional tiers for Qwen Cloud Token Plan Individual (sorted by third parties, subject to the official website):
| Tier | Monthly Promotional Price | 5-hour Credits Limit | 7-day Credits Limit |
|---|---|---|---|
| Lite | $6 | 700 | 2,500 |
| Standard | $18 | 3,000 | 10,000 |
| Pro | $68 | 12,000 | 40,000 |
The model consumption during the preview period is approximately 1/10 of the standard price; during some time periods (UTC+8 22:00–08:00), it can be as low as 1/50. Note: This is a Credits-based system, not the announced official API pricing for Qwen 3.8.
2. Qoder / QoderWork¶
- Qoder: Targeted at code repositories and software engineering scenarios;
- QoderWork: Targeted at knowledge workflows such as documents and spreadsheets.
The two share the preview quota logic with the Token Plan, suitable for developers who do not want to build their own Agent .harness.
3. Third-Party Agent Tool Access¶
Qwen Cloud’s documentation lists tools that can be linked with the Token Plan, including Codex, Claude Code, Cursor, OpenCode, Cline, etc. Key access points:
1. Select the model ID qwen3.8-max-preview when using the API Key;
2. Record the test date and reasoning level (default xhigh, optional low / high / xhigh);
3. Fix the prompt, tool permissions, and repository snapshot when comparing evaluations;
4. Count not only the answer quality, but also Credits consumption, latency, and retry counts — frequent inference can cause significant cost fluctuations for the same prompt.
4. Known Integration Layer Restrictions (Codex Documentation)¶
- Input: Text + images (video/document capabilities are subject to the latest official documentation);
- Parallel tool calls: Currently marked as unsupported in metadata;
- Built-in tools: Web search, code interpreter, crawler, image-to-image search, text-to-image search, etc. (Harness tools).
Developer Evaluation Checklist: Preview Can Be Tested, Not Suitable for Production Deployment¶
Combining the verified results above, it is recommended to adopt the following “graded citation” rules when making technical selections or publishing科普 articles on public accounts:
1. Can be cited as fact: Launched on July 19, 2.4T total parameters, MoE, multimodality, preview on three platforms, qwen3.8-max-preview ID, promise of open-weight soon;
2. Should be labeled as manufacturer’s claim: “Second only to Fable 5”, “frontier-level capability”;
3. Should be labeled as pending: Weight release date, license, active parameters, official API pricing, stable model behavior;
4. Should not be included in SLA for now: Any online business that relies on fixed behavior.
If you are choosing between Kimi K3 (with a July 27 weight release date) and Qwen 3.8, the end of July is a natural checkpoint: one side has a dated weight promise, while the other has a “soon” without a calendar. Claims with dates can be verified; claims without dates can only be waited for.
Conclusion¶
Qwen3.8-Max-Preview is not groundless — the 2.4T multimodal MoE, ultra-long context integration specifications, and real endpoints on the Token Plan all show that Alibaba is continuing to bet on its flagship models. But this launch also clearly exposes a new form of competition among Chinese frontier AI players: narrative can precede evidence, and openness can precede a calendar.
For developers, the most appropriate attitude right now is: actively trial the preview, strictly record configurations, cautiously cite rankings, and wait until weights and benchmarks are released before drawing conclusions. The day when an official Qwen 3.8 repository with a license appears on Hugging Face is when “soon” truly becomes a “date” — until then, it is just an unfulfilled option.