Z.ai: GLM 5.3 API

GLM-5.3-Flash is the first natively multimodal model in the GLM-5 series with 320B total parameters and 18B active parameters. It incorporates several architectural improvements over GLM-5.2 including a hybrid architecture, sharply reducing long-context serving costs while preserving precise long-context and Manifold-Constrained Hyper-Connections (mHC) to further improve scaling efficiency.

Frequently Asked Questions

What is the context window of GLM 5.3?

GLM 5.3 supports a context window of up to 1,000,000 tokens.

Does GLM 5.3 support function calling?

Yes. GLM 5.3 supports tool / function calling.

Does GLM 5.3 support reasoning?

Yes. GLM 5.3 is a reasoning-capable model.