DeepSeek Opens Core AI Operator Tools to Huawei Ascend; Industrial AI Leaders Map Scale-Up Path
DeepSeek released Ascend 950 versions of TileLang and five core AI libraries, matching Nvidia's operator toolchain, while Siemens and AQ Technology said industrial AI's next phase depends on reliable, repeatable deployment.
Operators determine how much of a chip's theoretical performance is usable. A poorly optimized operator can leave expensive accelerators waiting on data and deliver only about 20% utilization; well-tuned operators can push utilization above 95%, Leiphone reported. Matrix multiplication, or GEMM, for example, requires manual management of cache levels, thread blocks, data prefetching, register tiling and instruction reordering, often taking senior engineers weeks.
The old choices were hand-written CUDA, vendor libraries such as Nvidia cuBLAS, or high-level DSL compilers such as OpenAI's Triton. Hand-written CUDA offers the highest ceiling but low development efficiency and locks developers to Nvidia. Vendor libraries perform well for standard operators but are inflexible. Triton lowers the threshold but has a performance ceiling and weak native support for non-Nvidia hardware, including Ascend and AMD. TileLang takes a fourth path: developers declare where each data block should be computed, and the compiler handles data movement, thread synchronization and other low-level work. It focuses on GEMM, dequantized GEMM, FlashAttention, linear attention and sparse MLA, operators that carry more than 90% of large-model training and inference computation, according to Leiphone.
TileLang reuses Apache TVM's compilation infrastructure. It was initially incubated by a Peking University team on TVM and later adopted and productionized by DeepSeek, which donated it back to the open-source community. By September 2026, the repository had 7.6k stars and 1,992 commits. Leiphone reported that a TileLang-optimized sparse attention TopK operator for DeepSeek V3.2 ran 2 to 20 times faster than native PyTorch, and that its MLA, FlashAttention and LinearAttention implementations are commercially ready. On Nvidia H100, TileLang's core operators outperformed PyTorch and Triton and approached Nvidia's cuBLAS and official FlashAttention, while also running on AMD's MI300X near hardware limits, the report said.
TileLang's cross-platform design separates target backends from execution backends. Target backends define hardware-specific instruction syntax and optimization rules; execution backends handle general compilation, loading and launching. Each target backend follows a four-layer pipeline covering dialect translation, context, pass pipeline and host/device code generation. Nvidia CUDA is the main platform, while AMD ROCm, Apple Metal and Huawei Ascend 950 are officially supported; LLVM CPU, WebGPU and CuTe DSL are experimental, and Chinese chips such as MetaX, Moore Threads and Hygon are adapted through ecosystem projects, according to Leiphone.
DeepSeek's Ascend support is tied to a larger strategic bet. On Sept. 21, The Information reported that at a closed-door investor meeting, DeepSeek founder Liang Wenfeng described using Huawei or other domestic chips to train models as one of the company's 'biggest bets' and said it 'must succeed,' according to Leiphone. Bloomberg had reported that DeepSeek placed a $2.56 billion order for at least 160,000 Huawei Ascend 950DT accelerators for a gigawatt-scale data center under construction in Ulanqab, Inner Mongolia, with deliveries potentially starting in the fourth quarter of 2026. The Ascend 950DT chips are mainly for inference, while training still largely relies on Nvidia.
The migration faces several obstacles. DeepSeek V3 and V4 training code was written for Nvidia's Tensor Core and Hopper WGMMA instructions, so running on Ascend would require replacing cuBLAS and NCCL calls with Ascend C and HCCL, and validating self-developed operators such as FlashMLA, DeepEP and DeepGEMM on Ascend 950. Communication primitives are fundamental to parallel training; NCCL and HCCL differ in API design, performance curves and fault-tolerance models, and any mismatch in collective communication could silently invalidate a training run, Leiphone reported. Supply is another risk: the same U.S. export controls pushing DeepSeek toward domestic chips also constrain Huawei's ability to expand Ascend 950 production, and Huawei faces shortages of advanced memory and other components. DeepSeek is training a 2-trillion-parameter model, above the 1.6-trillion-parameter flagship V4, and Liang has said the company has planned an 8-trillion-parameter model. He estimated that training an OpenAI-level model would require about 50,000 Nvidia GB300 processors, equivalent to 200,000 Huawei Ascend 950 chips. DeepSeek still cannot fully shed its reliance on Nvidia hardware in the short term.
Separately, QbitAI hosted a dialogue with Gu Xin, Siemens Xcelerator China's head of sales and ecosystem success, and Huang Yao, founder and CEO of AQ Technology, on where industrial AI is heading. Huang began with semiconductor wafer dicing. At micro and nanoscale, incoming material differences, surface textures and model changes affect equipment performance, and customer requirements vary, often requiring engineers to handle warnings, downtime and parameter adjustments. AI first improves recognition accuracy and generalization; AQ Technology then uses multimodal models and agents to capture some equipment engineers' operating experience so the AI can move from detecting problems to analyzing and handling them. Huang said if recognition is inaccurate, the wafer can be cut wrong and the loss is very large. In the case he described, agent technology raised equipment OEE and capacity by about 25%.
Gu said industrial AI's key requirement is industrial-grade reliability. Unlike general AI, which can regenerate an inaccurate answer, a wrong judgment in a factory affects production efficiency. Industrial AI therefore needs to be trustworthy, traceable and quantifiable, he said. Both panelists described industrial applications as highly fragmented: different sectors, scenarios and processes, and even different plants within the same company can differ in equipment, data, workflows and standards. Industrial data is often scarce, hard to obtain and tied to process or commercial secrets, making internet-style scaling difficult. Gu said landing industrial AI is a system engineering problem, not just an algorithm or model problem; below it are equipment, data and compute, and above it are process, control systems, business workflows and delivery.
Huang said the first half of the market will be led by high-value applications, with broader growth later. The current opportunity lies where high value, data readiness and technical feasibility intersect, especially yield, process and OEE applications in high-end manufacturing such as semiconductors and PCBs, where a one- or two-percentage-point change can translate into substantial profit. Gu said improving semiconductor equipment OEE by a few points would create huge economic value for customers. Siemens Xcelerator aims to act as a connector and amplifier, linking technology to real customer needs and scaling validated solutions with integration and delivery partners. A Siemens Chengdu factory and AQ Technology project used a proof of concept to validate industrial AI vision for a high-value inspection step. Gu said '0 to 1 proves it can be implemented in the industry; 1 to 100 is commercial success.'
Editor's Summary
DeepSeek's release of Ascend-compatible TileLang and five core libraries marks a step toward a domestic alternative to Nvidia's CUDA-based operator ecosystem, though training migration and Huawei supply constraints remain. In industrial AI, Siemens and AQ Technology said scaling from single projects requires reliable, traceable systems and ecosystem partners that can repeat delivery across fragmented factories. The two developments point to a Chinese AI stack advancing in both core compute tools and factory-floor deployment.