Zhipu Launches GLM-5.3-FlashX: Speed Reaches 200 Tokens/s, Domestic Computing Power Advances Again
Zhipu AI launches GLM-5.3-FlashX, with the API also going live, offering a maximum output of 200 tokens/s. It focuses on intelligence, price, and speed, providing high-throughput, low-latency inference for enterprise developers. The predecessor, GLM-5.3-Flash, was previously introduced overseas under the name Ox Alpha. It gained popularity due to its strong intelligence and cost-effectiveness at the same size, with increasing usage volume. Zhipu is now supporting growing demand.