Recently, Zhipu officially launched and open-sourced the first native multimodal model of the GLM-5 series, GLM-5.3-Flash (320B-A18B), promoting cutting-edge artificial intelligence technology to be more accessible with leading performance and ultra-low pricing.
The model's capabilities have reached the global forefront. In the Artificial Analysis Intelligence Index evaluation, it scored 57 points, equaling Claude Opus4.8, with comparable programming abilities and superior comprehensive performance compared to the larger parameter model GLM-5.2. Before its official release, the model was tested publicly under the anonymous identity "Ox-Alpha" and topped the usage charts on OpenCode and OpenRouter platforms.

The cost advantage is a core highlight. The price of GLM-5.3-Flash is only one-tenth of the same series GLM-5.3, with a limited-time discount as low as one-twentieth, equivalent to one-fortieth the price of Opus4.8, significantly lowering the entry barrier for using cutting-edge models.
In terms of architecture, the model adopts a hybrid structure of sparse attention and linear attention, making it the first open-source cutting-edge model to implement this solution, combined with manifold constraint super connection technology. Compared to previous generations, the attention computation has decreased by 3.01 times, and the KV cache has been reduced by 4.44 times, balancing long context capabilities with inference costs. At the same time, it is natively integrated with visual capabilities, capable of independently completing tasks such as front-end development, 3D modeling, and document creation using visual feedback, showing strong performance in professional scenarios like financial reports and legal documents.

Notably, the model's full traffic is supported by domestic chip clusters. The team developed a self-researched inference engine, relying on the EPD separation architecture and multiple memory optimization solutions, achieving a threefold improvement in end-to-end inference performance, with hardware utilization efficiency comparable to mainstream NVIDIA GPUs, proving the feasibility of domestic computing power supporting large-scale cutting-edge model inference.
Currently, GLM-5.3-Flash is fully open-sourced, offering API interfaces, online experience channels, and has been launched on intelligent Agent platforms such as ZCode and AutoClaw. It also offers daily experience cards for ten thousand uses, open for global developers to use.
Online Experience
Z.ai: https://chat.z.ai
Zhipu Qingyan App: https://chatglm.cn

