AI Agent Hub
Back to models
Z.ai: GLM 5.3 FlashX logo

Z.ai: GLM 5.3 FlashX

Multimodal z-ai Released 2026-09-18
42.0 / 100 320.0B params 1M context Proprietary

About this model

GLM-5.3-FlashX is Z.ai's high-throughput serving tier for GLM-5.3-Flash, released on September 18, 2026. It runs the same 320-billion-parameter mixture-of-experts checkpoint with roughly 18 billion active parameters per token, native multimodal inputs (text, image, video, and files), up to one million tokens of context, and always-on extended reasoning with configurable reasoning effort. FlashX is optimized for peak generation speeds of about 200 tokens per second on Z.ai's infrastructure rather than a separate weight release or benchmark harness.

Benchmark Scores

BBH
86.6
HLE
55.3
MMLU
88.1
Aider
81.6
IFEval
72.0
DeepSWE
63.4
MMLU-Pro
84.7
SimpleQA
33.5
AIME-2026
92.7
Arena-Elo
1541.0
GDPval-AA
1773.0
HellaSwag
87.1
LiveBench
78.9
BrowseComp
75.9
GPQA-Diamond
90.1
Context-Arena
79.4
LiveCodeBench
37.6
NL2Repo-Bench
56.3
AutomationBench
48.8
IMO-AnswerBench
82.5
Agents-Last-Exam
26.3
SWE-Bench-Verified
77.8
Terminal-Bench-2.0
61.1
Terminal-Bench-2.1
84.3
Toolathlon-Verified
78.4

Technical Specs

  • Parameters: 320.0B
  • Architecture: MoE Transformer with hybrid sparse and linear attention
  • Context Window: 1,000,000 tokens
  • Input Modalities: text, image

Hardware Requirements

  • Compute: API only

Pricing

Input Output Currency
0.37 / 1M tokens 1.25 / 1M tokens USD