The "price war" and "capability ceiling" in the large model industry have recently seen a new breakthrough. The mysterious model called "Niu Lai" (Ox-Alpha), which had previously attracted widespread attention in open-source communities, has finally been unveiled. It is actually the latest release from Zhipu AI, GLM-5.3-Flash. As the first native multimodal model in the GLM-5 series, it not only competes head-to-head with global top-tier models in performance, but also rewrites the industry competition rules with its highly disruptive pricing and fully domestic computing power infrastructure.
Regarding core performance and pricing positioning, GLM-5.3-Flash demonstrates a leap in cost-effectiveness. Its Artificial Analysis Intelligence score reaches 57 points, surpassing the previous flagship GLM-5.2 and exceeding the official version of DeepSeek V4 Pro by 53 points, far exceeding the average of 18 points for similar models. In terms of price, it breaks the traditional rule that "the larger the parameters, the more expensive the model." The input cost is only 0.8 yuan per million tokens, output is 2.8 yuan, and cache hit is 0.23 yuan. The cost is about 1/40th of Claude Opus 4.8, allowing cutting-edge intelligence to truly escape the dilemma of "using it sparingly."
The comprehensive native integration of multimodal visual capabilities is the biggest technical highlight of this model. Traditional programming large models rely on a one-way process where "humans act as judges," while GLM-5.3-Flash integrates visual encoding capabilities into the coding loop. When writing code, designing games, or building 3D scenes, the model can actively "observe" rendering results and interaction feedback, forming a "generate-observe-revise" closed loop. Official tests show that the model can even run independently for 16 hours in Blender, building a professional chef's home and test kitchen of about 400 square meters from scratch, demonstrating strong autonomous visual judgment and iterative improvement capabilities.
More importantly, the computing power behind this cutting-edge large model has been fully transitioned to domestic chips. On platforms such as OpenCode and OpenRouter, the global real-time online traffic supported by clusters is powered by domestic chips, and through a series of radical optimization technologies such as the dedicated SGLang inference engine, ReplaySSM, hybrid cache quantization, and EPD separated architecture, it overcomes memory and bandwidth bottlenecks. This breakthrough proves that domestic chips not only can perfectly support cutting-edge large models, but can also maintain stability and economy under large-scale real traffic.
Currently, the model weights of GLM-5.3-Flash have been open-sourced on platforms such as Hugging Face under the MIT license, and the API is also open to developers worldwide, marking a historic step forward in the integration of high performance, low cost, and domestic compatibility in large models.



