BEIJING, Aug. 26, 2026— Chinese artificial intelligence company Zhipu has released and open-sourcedGLM-5.3-Flash, a new model designed to bring advanced AI capabilities to users at significantly lower inference costs.
According to Zhipu, GLM-5.3-Flash is the first native multimodal model in the GLM-5 family. The 320-billion-parameter model activates about 18 billion parameters and is designed for coding, visual understanding, agentic tasks and professional workflows.
A Focus on Cost-Efficient AI
Zhipu said GLM-5.3-Flash was built around a new architecture aimed at reducing computing requirements while maintaining strong model performance.
The company said the model achieved a score of 57 on the Artificial Analysis Intelligence Index, placing it within the frontier-model range. Zhipu also said its coding performance was comparable to Anthropic's Claude Opus 4.8 in its internal evaluation.
The company has positioned pricing as one of the model's major advantages. Zhipu said GLM-5.3-Flash is priced at roughly one-tenth of GLM-5.3, with a limited-time discount reducing the price further.
Native Multimodal Capabilities
Unlike earlier GLM models focused primarily on text and coding, GLM-5.3-Flash incorporates native visual capabilities.
Zhipu said the model can use visual feedback during coding tasks, allowing it to inspect interfaces, rendered results and interactive environments before making further changes. This is designed to improve tasks such as front-end development, game development and 3D applications.
The model can also work across code, browsers and graphical interfaces through agent systems, enabling it to inspect and verify its own output rather than relying solely on text-based reasoning.
Designed for Professional Work
Zhipu is also targeting applications beyond software development.
The company said GLM-5.3-Flash has been optimized for document-related tasks involving formats such as PPTX, PDF, DOCX and XLSX. Its visual capabilities allow the model to examine generated documents and improve their presentation.
Zhipu also highlighted financial and legal applications, including research based on sources, financial analysis, contract review and document drafting.
Built for Long-Context Applications
The model combines linear attention and sparse attention to reduce the computational cost of processing long contexts. Zhipu said this architecture can reduce attention computation and key-value cache requirements compared with GLM-5.3.
The company also said GLM-5.3-Flash supports a context length of up to one million tokens, making it suitable for applications involving large amounts of code, documents and other information.
Part of Zhipu's Rapid Model Development
The release follows Zhipu's launch of GLM-5.3 on Aug. 14. Zhipu described GLM-5.3 as a major upgrade focused on long-running agent tasks and coding capabilities. The company reported significant improvements on several software-engineering and agent benchmarks compared with GLM-5.2.

GLM-5.3-Flash takes a different approach by emphasizing efficiency, multimodal interaction and lower operating costs.
Zhipu said the new model is available globally through its API and online services, while its model weights have also been released through open-source channels. The company has additionally integrated the model into coding and agent products.
The release highlights a broader trend in the AI industry: competition is increasingly shifting from simply building larger models toward improving inference efficiency, multimodal capabilities and the cost of deploying advanced AI at scale.

微信扫一扫打赏
支付宝扫一扫打赏