Z.ai: GLM 5.3 FlashX
42.0 / 100
320.0B params
1M context
Proprietary
About this model
GLM-5.3-FlashX is Z.ai's high-throughput serving tier for GLM-5.3-Flash, released on September 18, 2026. It runs the same 320-billion-parameter mixture-of-experts checkpoint with roughly 18 billion active parameters per token, native multimodal inputs (text, image, video, and files), up to one million tokens of context, and always-on extended reasoning with configurable reasoning effort. FlashX is optimized for peak generation speeds of about 200 tokens per second on Z.ai's infrastructure rather than a separate weight release or benchmark harness.
Benchmark Scores
BBH
86.6
HLE
55.3
MMLU
88.1
Aider
81.6
IFEval
72.0
DeepSWE
63.4
MMLU-Pro
84.7
SimpleQA
33.5
AIME-2026
92.7
Arena-Elo
1541.0
GDPval-AA
1773.0
HellaSwag
87.1
LiveBench
78.9
BrowseComp
75.9
GPQA-Diamond
90.1
Context-Arena
79.4
LiveCodeBench
37.6
NL2Repo-Bench
56.3
AutomationBench
48.8
IMO-AnswerBench
82.5
Agents-Last-Exam
26.3
SWE-Bench-Verified
77.8
Terminal-Bench-2.0
61.1
Terminal-Bench-2.1
84.3
Toolathlon-Verified
78.4
Technical Specs
- Parameters: 320.0B
- Architecture: MoE Transformer with hybrid sparse and linear attention
- Context Window: 1,000,000 tokens
- Input Modalities: text, image
Hardware Requirements
- Compute: API only
Pricing
| Input | Output | Currency |
|---|---|---|
| 0.37 / 1M tokens | 1.25 / 1M tokens | USD |