DeepSeek Open-Sources Ascend Infrastructure Components With Huawei to Build AI Chip Software Ecosystem
On Sept. 30, DeepSeek released open-source infrastructure components for Huawei's Ascend platform, including TileLang tools, high-performance operator libraries and the DeepEP communication library, as Huawei detailed joint supernode networking and inference benchmarks.
DeepSeek released several high-performance Ascend operator library components, including DeepGEMM, FlashMLA, TileKernel and DeepSelect, along with the DeepEP distributed communication library. Huawei's Ascend platform provides stable, open Ascend C API interfaces that allow users to optimize memory access paths and compute pipelines for high-performance operator development. With the open Ascend C programming interface and the PTO ISA low-level instruction system, Ascend supported DeepSeek's open-source TileLang programming work. The Ascend ecosystem supports both fine-grained manual tuning by experienced developers and compilation from high-level languages such as TileLang.
Huawei provided DeepSeek with jointly defined Ascend SuperPoD Flex and UBL128 networking solutions. The design can achieve a 128-card 3.2 Tbps single-layer switched Scale-up network and a 256K-card two-layer switched Scale-out network, supporting ultra-low-latency inference and large-scale training of frontier foundation models. Based on the fully interconnected UBL128 supernode, Ascend provides the ASC-COMM high-performance custom communication programming library to support high-performance programming of user communication operators and communication-computation fusion operators. DeepSeek developed the high-performance DeepEP communication library, covering communication operators under EP, CP, PP and FSDP modes to support extending larger models to larger clusters. Measured interconnect bandwidth reached Dispatch 375 GB/s and Combine 347 GB/s, close to the hardware limit.
To support developers deploying DeepSeek models on Ascend 950 and supernode clusters, Huawei open-sourced the results of the joint innovation in the CANN community. These include large-EP low-latency inference deployment, single-card and single-machine deployment, large-scale training, ultra-long-text KVcache pooling and Agentic RL.
For large-scale inference scenarios, based on an EP32 deployment strategy in offline inference mode, the pure model performance of DeepSeek-V4.1-Flash without a framework can reach TPOT=5 ms with per-card output throughput of 2,469 tokens/s, and TPOT=10 ms with per-card output throughput of 5,102 tokens/s. The benchmark data were collected in offline inference mode and do not include the effects of serving scheduling and framework load balancing, with ContextLength set at 128K and a Dspark speculative acceptance rate of 0.85, according to the announcement.
Huawei said Ascend will continue to work with frontier model teams such as DeepSeek on chip-model co-design and joint innovation. The company said the aim is to use the AI computing base to carry model capabilities, adapt deeply in both directions, connect the full chain from inference to training, and build an open, efficient and easy-to-use AI software ecosystem. The related technical reports cover communication programming, operator development, inference optimization, training, quantization and agent deployment practices.