Huawei, Nvidia Target KV Cache Storage as Ruqi Builds Real-World Data Pipeline for Physical AI
Huawei and Nvidia target KV cache storage; Ruqi builds real-world data infrastructure for physical AI, Leiphone reported.
The core logic of external KV cache is 'storage for compute,' Zhang Heng, a storage scholar, told Leiphone. Historical KV already computed can be discarded and recomputed later or saved and reused when a later request hits. Zhang said if hit rate rises from 90% to 95%, recomputation falls from 10% to 5%, halving compute overhead for recomputation. Higher hit rates require keeping more historical KV; in the high-hit region, covering the last few percentage points of long-tail data may significantly increase storage capacity. Li Hui, head of a storage vendor, told Leiphone that data lifecycle varies. Consumer applications may clear KV cache after minutes or one or two hours. Some enterprise applications value context continuity and keep KV cache longer, making cross-request reuse more valuable.
Li said if DRAM were cheap and abundant, putting everything in DRAM would be simplest and fastest, but DRAM is expensive and scarce, leaving room for SSD. The replacement is not 1TB DRAM for 1TB SSD. In overall cost, DRAM and SSD are not a 1:1 substitution but '1:N,' sometimes 1:5 or 1:10. At the same capacity, DRAM and enterprise SSD costs can differ by dozens of times. That gap is pushing memory offload from a fallback for resource shortages into initial system design. Enterprises can replace some DRAM capacity with cheaper SSD, but need more SSD capacity and must still calculate capacity, performance and procurement cost. Zhang said most inference scenarios do not yet require independent KV storage devices; companies usually first exhaust existing DRAM, memory pools and local SSD before external storage. Huawei M900 and Nvidia CMX indicate that KV tiering, once mostly inside servers, is moving toward an independent shared storage layer.
Independent KV storage devices have appeared, but SSDs for KV Cache have not formed a mature, unified product category. Li said the industry is discussing SSDs optimized for AI inference and KV cache workloads, but most customers still buy ordinary enterprise SSDs. The problem is not that SSDs cannot be made, but that the workload KV Cache places on SSDs is not fully understood. Li said one related scheme initially required about 3 DWPD, then raised it to 9 DWPD; DWPD means drive writes per day. 'It might become 12 tomorrow and 5 the day after,' he said. The metrics depend on how long KV Cache is kept. If cache is kept for minutes, data is updated frequently and endurance pressure is higher. If data is kept for weeks, write frequency falls and capacity becomes more important. Li estimated that industry consensus on a more unified product may take six to 12 months; after standards are set, product development, validation and optimization may take about another year. North American customers have already increased SSD purchases because of AI inference and KV cache; China has also begun adding orders, but more slowly.
External KV cache directly relieves DRAM load. Zhang said the core goal is to swap memory, not video memory. Historical KV that once stayed in DRAM can sink to SSD, reducing memory configuration. The effect on HBM is less direct. Zhang said external storage mainly affects first-token latency: the speed of retrieving historical KV affects how soon the model starts generating the first token. Once generation starts, the relevant KV is back in VRAM, and later generation speed still depends on high-speed VRAM. Gu Huai, founder of an AI chip company, told Leiphone that compute-in-memory tries to reduce the movement of model weights and current data. In traditional inference, weights and intermediate data move frequently between storage and compute units. CIM aims to compute some tasks near the data. Gu said the advantage is especially clear in the Prefill phase, where the same model parameters can serve multiple requests under high concurrency. In the Decode phase, different requests use different KV caches, so less data can be reused, and CIM's benefit is less obvious than in Prefill. Gu said CIM could lower first-token latency and support higher concurrency if deployed in servers, but it cannot replace the need to store historical KV cache.
Ruqi Mobility began thinking in 2023 about what role a mobility platform could play in AI beyond providing Robotaxi and other travel services, Han said. It launched a smart-driving data toolchain that year, collecting driver behavior, road conditions and road network data from daily operations to support autonomous driving algorithm iteration. Over two years, Ruqi built a data closed loop connecting drivers, Robotaxis and city roads on one end and model training and iteration on the other. The rise of embodied intelligence gave the system a new use. Han said autonomous driving and intelligent robots do not need identical data. Autonomous driving mainly performs perception, decision-making and movement in two-dimensional road space. Robots enter three-dimensional space and handle grasping, carrying and fine operations. But the data processing flow is highly similar: collection, preprocessing, labeling, quality control and export. The underlying toolchain can be reused.
Real data collection is costly and slow, so more autonomous driving and embodied AI companies use world models and simulation platforms to generate synthetic data. Han said stronger world models do not reduce the importance of real data. Synthetic data can use real long-tail cases as a base to expand dangerous, critical and complex traffic scenarios that are hard to collect repeatedly. Real-world data continuously records unpredictable behavior by traffic participants, changes in road environments and physical responses from sensors, helping discover problems that have not been identified and providing calibration and validation for synthetic results. The two form a complementary loop: real data discovers and validates, synthetic data expands known problems.
Ruqi is trying to use data collection vehicles, Robotaxis and offline operations networks tied to routine mobility services to obtain real, continuous and scaled data at low cost, then convert it through its data loop into training materials for autonomous driving, world model and embodied intelligence companies. Han said customers need low-cost, high-quality data that can enter training processes directly, plus deterministic delivery capacity and engineering efficiency. Even a customer with strong algorithms and training capabilities needs continuous investment in devices, people and operations to conduct large-scale collection, compliance governance, labeling and quality control. Han described a project with a leading autonomous driving company founded in 2016. Its R&D investment and demand for data labeling grew quickly over two years. Ruqi structured complex labeling tasks according to the customer's data specifications, matched resources to different data processing tasks, managed the process digitally and delivered daily. It also intervened early in labeling work to shorten the ramp-up period when demand spikes. The customer received stable data capacity that could keep up with model iteration.
Han said the business does not have a high single-point academic technical barrier. The real barrier is long-term accumulation: how to collect data using sensors and onboard devices, how to deploy them and how to capture drivers' operations in real scenarios. The second layer is the data processing toolchain, from collection, cleaning, labeling and evaluation to format adaptation and customer delivery. Different customers have different data formats, sensor standards and model training requirements. The third layer is that capabilities change with customer feedback, which in turn changes collection methods, processing algorithms and toolchains. Han said data labeling is currently the most mature and highly automated link. Adapting a new customer may take more than a year to become standardized and automated, but accumulated tools and processes now shorten that cycle. Ruqi does not want headcount to grow in proportion to business growth. It aims to extract repeated capabilities into automation tools and standardized platform capabilities. Repeat customers already account for a certain proportion, and repurchase rate is a core internal KPI, though Han did not disclose the number.
Han said Robotaxi data can only be partially reused for embodied intelligence. Much autonomous driving data is perception data, which can support environmental perception training for robots. Driving is a relatively unified task, while embodied intelligence entering logistics may require moving objects, entering supermarkets may require arranging shelves, and entering production lines involves fine operations. It needs robotic arm operation, motion trajectories and fine action data, which require new sources. The reusable part is the data infrastructure: data import, preprocessing, cleaning, quality control, export and compliance governance. For embodied intelligence, Ruqi needs to add capabilities for three-dimensional space and fine action labeling, but the overall framework does not need to be rebuilt from scratch. Ruqi chose car aftermarket services as its first embodied intelligence scenario. Han said the company looks for real business demand, the ability to continuously organize data collection and synergy with existing business. Ruqi has car services for traditional mobility and is building an offline operations network around Robotaxi. Interior and exterior cleaning, appearance inspection and charging plug insertion and removal are real operations and tasks embodied robots may take. In June 2026, Ruqi launched an embodied intelligence data platform that processes ego-centric first-person operation video, including data import, AI preprocessing, action labeling, multi-level quality control and standardized format export.
Han said it is hard to give a unified ratio for how much daily data is valuable because validity is relative. A company that has been in autonomous driving for years may already have simple, common road scenes and now needs emergency braking, complex roads and extreme long-tail scenarios. Another company at a different stage may need different data. Some data is clearly useless: LiDAR hitting positions 10 or 20 stories high is not useful for ground autonomous driving and is removed during cleaning. Customers look at quality, authenticity, scenario scarcity and whether data can be generated continuously at scale. In embodied intelligence, the threshold for real scenarios may be higher. A robot company entering a real industrial production line to collect data is not easy; manufacturers consider production safety and data leakage, and do not casually let external robots and teams into production lines. Han said the data most likely to be replaced or devalued is routine data collected repeatedly to increase volume, not real data itself. Synthetic data is suited to expanding known problems. But the real world also discovers problems that were not known, and these cannot be fully generated by a model and then verified by the model itself. Han said he hopes that in three years, outsiders will still see Ruqi Mobility as a mobility service company, but also as important infrastructure in the AI data ecosystem.