Qualcomm Unveils 6th-Gen Snapdragon 8 Super Ultimate for Agentic AI; Zhongke Leinao Builds Token Factory
Qualcomm launched two 2nm mobile platforms for agentic AI, led by a 5GHz Snapdragon 8 Super Ultimate with on-device 30B MoE support. Zhongke Leinao unveiled a compute-electricity-Token platform and reported a cross-region scheduling test.
Qualcomm launched the sixth-generation Snapdragon 8 Super Ultimate and the sixth-generation Snapdragon 8 Ultimate at the same time. Chris Patrick, senior vice president and general manager of mobile handsets at Qualcomm Technologies, said a single standard can no longer define a flagship, and that maintaining leadership requires more than one path forward. The two platforms share core technologies, with the Super Ultimate at the top of the portfolio. It includes what Qualcomm calls the industry's first mobile CPU with a maximum clock speed of 5GHz, Matrix Cores added to the GPU for the first time, and an NPU capable of running a 30-billion-parameter MoE model. Zhu Yuankun, a Qualcomm product technology expert, told Leiphone that the main problem for on-device agents is memory and battery life, not performance alone.
The memory wall challenge runs across the CPU, GPU and NPU. The 5GHz Oryon CPU adds Oryon FlexCache, allowing eight cores to access a shared L2 cache pool and dynamically allocate resources based on workload, so data does not have to be reloaded from system memory at every task handoff. The Adreno GPU has a second-generation 18MB high-speed memory module, HPM, to hold rendering tiles, frame buffers and intermediate results, reducing repeated access to system memory and cutting power by 12%. The Hexagon NPU's shared memory increases by 50%, keeping more model state, context and KV-cache on-chip and reducing external DDR access. The platform also uses LPDDR6 memory and UFS 5.0 flash. On the NPU, Qualcomm strengthened Transformer support with the Hexagon Element Accelerator, enabling up to 32K Token context and up to 80% better prefill performance on some models. The sensor hub's dual Micro NPUs improve performance by 85% while reducing power by 20%. The Adreno GPU delivers up to 44% higher performance and 40% better energy efficiency, which Qualcomm says is its largest GPU generational jump, and it is the first time Matrix Cores have been added to Adreno.
Qualcomm worked with Jieyue, Wulianghuo and Longsys to deploy StepEdge-Omni 30B-MoE on handsets. The 30-billion-parameter model activates about 3 billion routing parameters per generated Token. Through inference engines and heterogeneous scheduling across CPU, GPU and NPU, the model's runtime memory requirement is more than 50% lower than conventional approaches, with prefill speed above 330 Token/s and decode speed above 28 Token/s, according to Leiphone. It can complete continuous tasks such as email understanding, itinerary planning, calendar synchronization, flight and hotel recommendations, and email drafting. The dual-engine NPU and GPU prefill throughput is more than 30% higher than a general inference scheme using only the NPU. The second-generation AI-ISP architecture doubles video throughput and supports 8K 60fps video capture on a Snapdragon mobile platform for the first time. Qualcomm is extending its computing portfolio from phones to PCs, cars and other agent terminals, while linking to cloud computing resources for hybrid AI.
Separately, according to QbitAI, Zhongke Leinao, a company founded nine years ago, is positioning itself as a compute-electricity-Token factory. It launched a compute-electricity-Token integrated platform at the World Manufacturing Convention, following a 1+3 architecture: one decision brain plus compute, Token and electricity subsystems. Its BitaHub cloud platform manages Nvidia, Ascend and other chip architectures, aggregates model APIs and Token services, and connects enterprise private compute pools with public resources. The company says BitaHub has connected more than 3,000 computing nodes, with total scale exceeding 5,000P, and has served more than 80,000 enterprise and research users. Chang Feng, partner and deputy general manager, said industrial electricity usually costs a few tenths of a yuan per kWh, becomes worth tens of yuan as compute, and can exceed 100 yuan per million Token for some models. Customers, he said, ultimately buy Token services that meet task and metric requirements.
Ding Haisong, CTO and president of the research institute, described the system as monitoring, forecasting and decision-making, with decisions bounded by user experience, result quality and system stability. It identifies flexible loads in computing tasks: real-time tasks sensitive to first-Token latency stay in place, while batch tasks that can be queued or interrupted can move to low-price periods and locations. Heterogeneous chips are managed through a foreman-style plug-in for each chip type, with capability profiles recording what each chip does well. The system can also split Prefill and Decode in large-model inference and connect the stages through KV Cache migration, sending compute-intensive Prefill to stronger chips and memory-bound Decode to chips with large memory and high throughput. In a cross-region test spanning Shanghai, Wuhu and Urumqi, control response stayed within 200 seconds, cross-region migration succeeded 100% of the time and energy forecast accuracy reached 98%. Zhongke Leinao called it China's first cross-three-city compute-electricity collaborative intelligent scheduling test. On Sept. 17, the Wanjiang Artificial Intelligence Science and Technology Industrial Park opened in Urumqi, with its first integrated computing center exceeding 1,000P already in operation.
The company is targeting a market in which Token demand is rising while prices fall. National Data Administration data cited by QbitAI show daily Token calls in China rose from 100 billion in early 2024 to 140 trillion in March 2026, more than a 1,000-fold increase in two years, while some mainstream model APIs have cut prices cumulatively by more than 90%. Zhongke Leinao uses effective Token per kWh as its top-level optimization metric, with three layers: compute and Token optimization through heterogeneous scheduling and inference acceleration, compute-electricity synergy to lower unit Token costs, and organizing general Tokens into scenario results priced by value. It has formed commercial loops in AI for Science and energy. Financing has followed: China Mobile Fund invested exclusively in its Series B in 2025, CRRC Capital led a several-hundred-million-yuan Series B+ in 2026, and Ginkgo Valley, Shuimu and Qidi participated.
Editor's Summary
Qualcomm's two 2nm mobile platforms and its on-device 30B MoE deployment mark a shift in flagship mobile chips toward agentic AI workloads, with memory, scheduling and power efficiency as key battlegrounds. Zhongke Leinao is pushing a compute-electricity-Token operating layer, backed by a cross-region scheduling test and a new industrial park, as Token prices fall and system efficiency becomes the focus of competition. The outcomes will depend on real device experience and large-scale operations.