AI News Feed
Market watch
AI Chips & Compute

GMIF 2026: 140 Trillion Daily Token Calls Push SSDs Into New AI Roles

At GMIF 2026, executives said China's daily token calls exceeded 140 trillion as of March, forcing SSDs, NAND and data tiering to take on new roles in AI inference. The summit detailed QLC SSD replacement of HDDs, Z-NAND and HBF efforts, and active scheduling of KV cache and MoE weights.

Morgan Stanley executive director and Greater China semiconductor analyst Yan Zhitian said that excluding HBM, memory would account for nearly 50% of global cloud data center capital expenditure in 2026 and is expected to rise to 53% in 2027. Samsung Electronics vice president and CTO of the Memory Business Kevin Yoon converted the 140 trillion figure into a comparison: China's daily token consumption is more than 20 times the token volume of all books in the National Library of China, and close to six times the total token volume of all published books in human history.

The pressure on storage comes in three forms, according to Solidigm Asia-Pacific sales vice president Ni Jinfeng: data richness, inference complexity and operational persistence. He said that when storage costs were lower, the industry tended to simplify storage tiers, but rapid growth in AI data demand and rising cost pressure are further subdividing the original storage pyramid.

Xinghe Zhilian demonstrated an in-vehicle Agent that processes voice, vision, vehicle signals, dialogue history, long-term memory, user profiles and real-time vehicle status. Arm global storage business head John Xavier Lionel said larger models create more weights and runtime state; longer context increases the KV cache needed for subsequent token generation; more users require separate KV cache for each request; and the decode stage repeatedly accesses model weights and KV cache. Dify co-founder Chen Lusha predicted that enterprises may run hundreds of Agents, rather than a few, on 24-hour duty for long-horizon tasks and high-concurrency requests.

Parallel Tech's MaaS model lets enterprises call different models through APIs and pay by token, requiring providers to maintain first-token latency, success rates and throughput while monitoring call costs. For building local token factories, Lenovo focuses on actual utilization of compute resources, stability of service operations and the cost of continuously producing tokens rather than peak performance.

SSD vendors are upgrading on several fronts. Silicon Motion uses host software and the storage software stack to identify task types, after which the SSD controller adjusts resource allocation by QoS level so that one SSD can respond to different AI loads. SanDisk senior product marketing director Zhang Dan divided SSDs into three types: SSDs directly connected to GPUs for hot KV and other frequently accessed data; network-connected SSDs that still serve compute but can be deployed at larger scale; and storage SSDs outside the compute cluster that emphasize capacity for data not frequently accessed.

Ni Jinfeng observed that large North American internet companies have begun replacing some HDDs with high-capacity QLC SSDs to reduce floor space and power consumption. HDDs are not disappearing. Dashu Jixian's PMD, positioned between traditional HDDs and enterprise SSDs, uses multiple parallel channels, larger I/O request queues, request aggregation, prefetching and caching to improve storage bandwidth and random access. Oconnor added deep learning I/O loads and scenarios such as KV cache and checkpoint recovery to AI SSD testing, because AI workloads produce bursty, highly concurrent I/O behavior rather than the fixed 4K and 128K block reads and writes used in traditional SSD tests.

NAND is also breaking old boundaries. Samsung's Z-NAND uses SLC core dies and TSV technology to improve read performance; in Kevin Yoon's architecture, the operating system and KV cache remain in DRAM while large, read-intensive model weights are placed in Z-NAND. Another path upgrades NAND into HBF. Peking University School of Integrated Circuits dean Cai Yimao said there are two main technical routes: one shrinks the NAND array to reduce read latency, and the other borrows HBM's parallelism approach by increasing the number of interfaces to raise bandwidth. He said both still face problems, including density loss from array scaling, relatively high NAND read latency and limited write endurance for frequently updated data such as KV cache. Cai also said Peking University and partners including Yanxinwei are exploring the use of NAND arrays to perform some in-memory computing and reduce data movement between storage and compute units.

Data tiering is shifting from fallback to active scheduling. The traditional storage hierarchy balanced speed, capacity and cost, with different data types mapped to relatively fixed media. AI creates more data types, and the access frequency and lifecycle of the same data change during computation. Speakers said cross-layer data movement is not new in the AI era, but management is moving from reactive tiering, in which HBM insufficiency pushes data to DRAM and DRAM insufficiency pushes data to SSD, to active tiering at system design and dynamic scheduling during operation.

Xinzhan Speed vice president Li Zhen described the AI data supply system through three keywords: storage decides which medium carries which data, connectivity ensures stable and efficient data movement, and solutions handle scheduling and application adaptation. Lianyun and Yinpu's joint AI SSD scheme applies this division during inference. For MoE model parameters, expert weights needed at the moment remain in CPU memory and GPU memory, while temporarily unused weights are placed in SSD. Because SSD access is slower, the system predicts which experts may be called next and prefetches the corresponding models and parameters. For KV cache, high-frequency access remains in memory while low-frequency portions sink to SSD. Jiechuang Intelligent's Changqing Cloud platform prepares data needed for tasks on the appropriate nodes and matches data with suitable compute resources based on task type, GPU generation and network bandwidth. Baiwei Storage chairman Sun Chengsi said data supply efficiency determines compute output. As AI inference moves toward scale, high concurrency and continuous operation, storage must answer not only where data is placed but how to get it to compute in time.