Z.ai: GLM 5.3 API
GLM-5.3-Flash is the first natively multimodal model in the GLM-5 series with 320B total parameters and 18B active parameters. It incorporates several architectural improvements over GLM-5.2 including a hybrid architecture, sharply reducing long-context serving costs while preserving precise long-context and Manifold-Constrained Hyper-Connections (mHC) to further improve scaling efficiency.
- Context window: 1,000,000 tokens
- Max output: 128,000 tokens
- Input: text
- Output: text
- Reasoning: Supported
- Tool calling: Supported
- Released: 2026-08-25
Frequently Asked Questions
What is the context window of GLM 5.3?
GLM 5.3 supports a context window of up to 1,000,000 tokens.
Does GLM 5.3 support function calling?
Yes. GLM 5.3 supports tool / function calling.
Does GLM 5.3 support reasoning?
Yes. GLM 5.3 is a reasoning-capable model.