Damao Raises About RMB 200M as SenseTime and Huizhi Push AI Infrastructure Moves
On Sept. 23, QbitAI reported that Damao Technology closed about RMB 200 million in new funding, SenseTime said daily token service reached 4.5 trillion and detailed heterogeneous inference work, and Huizhi Intelligence launched Hellome, a platform designed to compress enterprise AI delivery to weeks.
Damao Technology, founded in 2021 and headquartered at Xuhui MOSPACE in Shanghai, is described in the report as the first Chinese AI company to scale energy large models in compute-power coordination, virtual power plants and broader energy scenarios. The new round follows an A+ round of nearly RMB 100 million in October last year led by Puquan Capital, an investment arm of CATL. The company builds compute-power coordination platforms, AI virtual power plant platforms and intelligent operations and maintenance platforms on self-developed vertical energy models and an AI agent cluster. It serves intelligent computing center customers across consulting and planning, investment calculation, intelligent operations, refined O&M and energy value-added trading.
Its core product, the Compute-Power Coordination 2.0 Platform, also called the AIDC green power direct-connect energy operation operating system, is designed to provide full-cycle smart energy operation for GW-level AI data centers using green power direct connection. Damao says the system can turn AIDC facilities from conventional electricity loads into dispatchable and tradable flexible resources under a new power system. Damao's shareholders include CATL, SenseTime and Cambricon. Compute-power coordination was written into China's government work report for the first time in early 2026. In September, the State Council called for advancing compute-power coordination and compute-network integration and accelerating green power direct connection and source-grid-load-storage projects. Industry analysis cited by QbitAI said hardware players are numerous in the sector, but companies that can provide complete underlying software and commercialize the full compute-power coordination chain are rare. Damao CEO Jian Yumin said in a previous media interview that the company positions itself as a core industry operation service provider rather than only a technical service provider. Its model combines technology enablement with operations, sells on value and shares in cost reductions and revenue gains, he said. Jian said the sector is in a fast window and the company will keep investing in research and market expansion.
In a separate QbitAI report, SenseTime's large-scale AI infrastructure unit detailed the system challenges of scaled inference at the 2026 Global AI Chip Summit's Token Factory heterogeneous mixed training and inference technology seminar. Luo Wei, a senior technical expert at the unit, said Agent-era workloads differ from single-turn dialogue: long context increases input computation and KV Cache usage, multi-turn conversations constantly change prefix reuse and hotspots, and bursty and long-tail demand keeps traffic and length distributions fluctuating. The infrastructure design has shifted to center on tokens, state and latency budgets, organizing resources around token production efficiency, delivery quality and unit cost. SenseTime's average daily token service volume grew from 0.46 trillion in February to 4.5 trillion in August, according to the report.
To turn single-point efficiency into system-level token capacity, SenseTime pursues horizontal orchestration and vertical optimization, Luo said. Horizontally, it builds unified resource profiles for chips of different brands and specifications, covering compute throughput, KV Cache capacity, effective bandwidth and model compatibility. Inference requests are managed across their lifecycle, from admission pairing and Prefill execution to KV Cache state handoff and Decode takeover. The platform dynamically adjusts chip roles rather than permanently fixing certain chips as Prefill or Decode nodes, Luo said, citing model, batch and context-length bottlenecks. Vertically, SenseTime optimizes the model, engine and chip layers, including weight quantization and memory budgeting, parallel optimization of Attention and GEMM operators, and compilation, memory access scheduling and memory management. The results are validated under real workloads.
SenseTime is moving from fixed Prefill/Decode ratios to dynamic resource pools. It has built a Prefill Pool and a Decode Pool, with request routing and P/D resource matching based on real-time load, KV Cache state and node health. When long input raises Prefill pressure, the system can add P nodes; when long output raises Decode load, it can add D nodes, and nodes can switch roles. For future multimodal and MoE models, the next step is module-level resource pools: splitting Encoder, Attention, FFN and visual encoding into independent pools that scale elastically. Luo said SenseTime will continue product and technology iteration with ecosystem partners to accelerate AI infrastructure upgrades and inclusive AI.
Also on Sept. 23, Nanjing-based Huizhi Intelligence released Hellome, which it called China's first FDE direct-connect application agent service platform. Clients submit intelligent needs on the platform; certified FDEs take on projects and use unified agent tools to design, deploy and launch solutions, compressing AI implementation to weeks, according to QbitAI. The company said demand has surged as large-model capabilities iterate, with organizations seeking AI for official documents, reports, public opinion monitoring and production management. The main bottleneck is delivery: generic AI products do not fit specific business processes, traditional custom projects are expensive and take quarters, and most organizations cannot build their own AI teams. Hellome is intended to provide a scalable and predictable delivery path.
Hellome platformizes the delivery process. After a client posts a requirement, a certified FDE takes the project and works with Hzhermes, Huizhi's desktop agent. Hzhermes runs in the local computer environment and can operate local software, process documents and spreadsheets, and call dozens of tools including browsers and search. One FDE handles requirements, solution design, deployment and later iteration, with no subcontracting. Clients can use mature agents packaged by FDEs or request customization. Because the agent developer and later service provider are the same person, response chains are shorter and responsibility clearer, the company said. Solutions developed in projects accumulate as skill packages covering dozens of common scenarios, including official document writing, data analysis, public opinion monitoring and content creation. Later clients can enable these packages directly, reducing delivery costs as cases accumulate.
At launch, Hellome already offered dozens of agents for official document writing, data analysis, public opinion briefs, livestream scripts, short video creation, novel writing and recruitment. For personalized needs, FDEs make incremental modifications on existing assets, with delivery measured in days rather than months. Huizhi also started an FDE Million Incentive Plan to lower entry and growth barriers for engineers through cash incentives, traffic support and skill certification. The company said many AI engineers are joining daily. Huizhi's agent services cover more than 50 industries and over 10,000 enterprise clients and have reached tens of millions of individual users. Its service, engine and compute product system includes Agentsyun Token Factory, which aggregates more than 200 large models for scheduling. Huizhi plans for Hellome to gather 100,000 FDEs in three years.