AI News Feed
Market watch
Robotics

Humanoid robot deployment needs patience, Zhejiang professor says at WRC

At the 2026 World Robot Conference, Professor Xiong Rong said humanoid robots still face reliability, generalization and precision gaps before becoming true productivity. Her team has secured 100-unit orders in 3C and 2,000 in apparel.

The report detailed that the team has signed orders at the level of one hundred units in the 3C field and 2,000 units in the apparel industry. The robots can perform car assembly, grasp flexible fabrics like shirt and denim cloth, operate template machines for pocket attachment, dispense transparent liquids with an accuracy of 1 milliliter, and handle logistics tasks such as stacking and unstacking, as well as stripping and laminating flexible screens. In terms of efficiency, a 5-meter transport task takes 35 seconds, and a full pocket-attachment process takes 50 seconds, reaching 80 percent of manual speed.

Xiong identified four gaps between demos and production lines. First, robots may fail in factories due to sudden conditions and abnormal situations, forcing new data collection and retraining. Second, generalization remains insufficient; current systems often require one-shot training for each new object, scene, or lighting change. Third, current operation precision is at the centimeter-to-millimeter level, while industry often demands sub-millimeter precision, which is beyond visual perception and needs force-touch feedback. Fourth, execution efficiency is low because the perception-reasoning-decision-control loop adds many steps compared to a pre-programmed motion. She stressed that the physical world is 'irreducible,' and neither expert models nor pure data accumulation can capture its full complexity.

To overcome these gaps, Xiong proposed breakthroughs in four layers. At the representation layer, she suggested fusing vision with force-touch perception to express the essence of physical interaction, which is difficult to represent with vision alone. At the perception layer, she called for enhancing spatial understanding in vision-language models and building prediction models for objects in both contact and non-contact states, forming a world model that evolves with embodied models. At the model layer, she emphasized the need for abstraction and analogical reasoning, as current methods are heavily dependent on data stacking. At the data layer, she highlighted low-cost collection of large-scale high-quality data through a 'Real to Sim to Real' flywheel.

Xiong also discussed ecosystem cooperation. Her innovation center works with upstream suppliers, downstream industry partners, data partners, and application developers. To address the shortage of talent in embodied intelligence, it has trained more than 200 teachers and 1,000 students by partnering with research universities, application-oriented colleges, and vocational schools, setting up curriculum systems and training bases.