inclusionAI: Ling 3.1 Flash
About this model
Ling 3.1 Flash is a hybrid-reasoning mixture-of-experts language model from InclusionAI (Ant Group). It scales to roughly 560 billion total parameters with about 25 billion activated per token, making it substantially larger than Ling 3.0 Flash while targeting flash-tier latency and cost. The architecture is tuned for agentic workflows, multi-step analysis, software engineering, search, and long-document tasks, with additional domain emphasis on medicine, finance, and materials science.
The model launched in late September 2026 as an API-only preview: weights were not publicly released at launch, though InclusionAI stated it intends to open-source checkpoints after the initial trial. Providers such as OpenRouter and Vercel AI Gateway expose a 262K-token context window during the promotion, while the design target is up to one million tokens once full service rolls out. It supports tool calling and extended reasoning modes suited to coding agents and terminal-style automation.
On vendor-reported and independently run evaluations, Ling 3.1 Flash ranks strongly among flash-class models on agentic and knowledge-work benchmarks, including high scores on BrowseComp, SWE-Bench Pro, Terminal-Bench 2.1, CyberGym, and GDPval-AA (Elo), with competitive Humanity's Last Exam results relative to other flash-tier systems. Treat vendor figures as directional until broader third-party replication is available.
Benchmark Scores
Technical Specs
- Parameters: 560.0B
- Architecture: Mixture-of-Experts Transformer
- Context Window: 1,000,000 tokens
- Input Modalities: text
Hardware Requirements
- Compute: API only
Pricing
| Input | Output | Currency |
|---|---|---|
| — | — | USD |