Z.ai: GLM 5.3 (batch)
About this model
Z.ai: GLM 5.3 (batch) is the discounted batch inference endpoint for GLM-5.3, Z.ai's flagship reasoning model aimed at complex software engineering, terminal agents, and long-horizon tool use. It shares the same 744B-parameter MoE backbone as GLM-5.2 (about 40B active parameters per token) with gains delivered entirely through extended post-training rather than a new pretrain. The API variant exposes a 1M-token context window, up to 128K output tokens per response, mandatory chain-of-thought via thinking mode, and configurable reasoning_effort levels (low, high, max; max is the default used in published benchmarks).
On public evaluations at max reasoning effort, GLM-5.3 ranks among the strongest open-weight coding agents: it reaches 88.2% on Terminal-Bench 2.1, 28.3% on Terminal-Bench 3.0, 66.9% on DeepSWE v1.1, and 84.5% on CyberGym, with emergent cybersecurity gains on exploitation-style suites. It also scores 62.5% on Humanity's Last Exam with tools, 73.0% on Toolathlon Verified, and an Artificial Analysis GDPval-AA v2 Elo of 1769. Independent harnesses (for example Vals AI and Mercor) report strong MMLU-Pro, GPQA-Diamond, LiveCodeBench, and SWE-bench Verified numbers, though scores vary by agent harness and effort setting.
Open weights are published on Hugging Face (zai-org/GLM-5.3, FP8 at roughly 756 GB), while the batch SKU is consumed through Z.ai and partner routers such as OpenRouter at reduced per-token pricing for asynchronous workloads. Typical use cases include coding agents (Claude Code-style harnesses), RAG over large repos, security research pipelines, and multi-step automation where cached context and batch pricing improve unit economics.
Benchmark Scores
Technical Specs
- Parameters: 744.0B
- Architecture: MoE Transformer
- Context Window: 1,048,576 tokens
- Input Modalities: text
Hardware Requirements
- Compute: API only
Pricing
| Input | Output | Currency |
|---|---|---|
| 0.45 / 1M tokens | 2.00 / 1M tokens | USD |