AI News Feed
Market watch
AI Chips & Compute

Zhipu's GLM-5.3-Flash Goes Live on SenseTime's Domestic Chip Infrastructure

Zhipu launched open-source GLM-5.3-Flash, running on SenseTime's domestic AI infrastructure, scoring 57 in A2I.

The model's architecture is designed for extremely low cost, combining Zhipu's latest 30 trillion-token multimodal pre-training corpus to achieve stronger performance with fewer computing resources. SenseTime's large-scale AI infrastructure supported the launch by providing large-scale domestic computing power and token services through an advanced heterogeneous mixed-inference technology, marking what the report calls a benchmark for scaled commercial use of domestic computing power.

SenseTime's support leverages a core capability system it calls “one model, one token factory, one agent management system.” The token factory organizes different chips, clusters, energy sources, and delivery nodes to continuously produce higher quality and lower cost tokens, forming a two-way feedback loop with the large model that adapts to iterative development needs. According to the report, this closed-loop capability enabled efficient delivery of the GLM-5.3-Flash launch.

Before the official release, GLM-5.3-Flash was tested anonymously as Ox-Alpha on OpenCode and OpenRouter, with token calls reaching 62 trillion, all executed on domestic chips. Compared with the initial baseline in the same hardware environment, the model's end-to-end service performance improved threefold, and hardware efficiency and single-token cost reached levels comparable to mainstream Nvidia GPUs. The report states this demonstrates that domestic chips can effectively and economically support frontier large-model inference in large-scale real business scenarios.

SenseTime has promoted domestic computing power from “single-point availability” to “large-scale commercialization.” At this year's WAIC, SenseTime addressed key bottlenecks in domestic computing power adoption with systematic solutions across cost-performance, adaptability, and energy efficiency. The report says its heterogeneous mixed-inference technology lets different domestic chip architectures play to their strengths, achieving overall inference cost-performance 1.25 times that of Nvidia H series and expanding token output about 2.5 times at equal cost compared with domestic homogeneous inference.

SenseTime also cut adoption barriers through low-level operator optimization, multi-card parallel tuning, and toolchain adaptation, and introduced the TPW (Tokens Per Watt) metric as a new value standard for AIDC. Its token factory delivered an average daily token service volume of 2.42 trillion in July, with a projection to exceed 10 trillion by the end of 2026 and a 25-fold annual increase. SenseTime said it will continue investing in token services based on domestic computing power to build high-performance, low-cost, and reliably deliverable AI infrastructure.