Embodied Intelligence: New Startup, World Models and Token Economics Take Center Stage
New embodied AI startup Lingxi Zhichuang launches; WRC 2026 debates world models and 'physical token' business models.
Lingxi Zhichuang unveiled its first-generation industrial embodied Harness system, ROSS, built on a 'model + harness' architecture. The system is named after cybernetics founder W. Ross Ashby. Rather than a simple engineering patch on top of a model, ROSS is designed as an industrial-grade execution system that converts a model's probabilistic action generation into stable, controllable and reusable physical-world execution. It features five core capabilities: model abstraction, skill abstraction, Agentic AI planning, layered safety monitoring, and a data flywheel.
Model abstraction standardizes inputs and outputs so different VLA or world-action models can be plugged into a unified interface, allowing new models to be integrated with partial adaptation while preserving upper-layer task planning, skill orchestration and safety policies. Skill abstraction packages tools, model capabilities and engineering strategies into reusable task units, turning scattered on-site know-how into standardized digital assets. The company also introduced layered monitoring and fault tolerance, with real-time safety handled at the controller level, trajectory monitoring in the middle layer, and task-level judgment carried out by agentic planning. In case of anomalies, the system can interrupt, retry, roll back, switch to alternative skills, re-plan, or ask for human takeover.
Duan, who studied at the University of Science and Technology of China under Professor Ji Jianmin, has invited Ji to join Lingxi Zhichuang as co-founder and chief scientist, forming a structure that combines academic guidance with engineering development. The ROSS Harness system is scheduled to appear at the upcoming second World Humanoid Robot Games, where it will participate in industrial scenario competitions involving continuous, complex and perturbed environments.
At the World Robot Conference (WRC) 2026, Chen Jianyu, founder and CEO of Xingdong Jiyuan, argued that VLA models are not the final answer. Chen said VLA's core is imitation—learning how humans act—which limits generalization to tasks never seen before. He contended that world models, which learn the rules of the physical world and predict future states, are likely to become the core paradigm for next-generation embodied intelligence. Xingdong Jiyuan has developed world action models that integrate video prediction and action prediction, claiming the architecture enables zero-shot task generalization and precise manipulation. The company has demonstrated tasks such as folding paper boxes, turning socks inside out, and using tools with a dexterous hand.
Also at WRC 2026, Gao Jiyang, co-founder of Xinghaitu, predicted that the sector's business model will eventually shift from selling robot hardware to selling 'physical world tokens'. He compared it to the way OpenAI charges for text tokens. In the final stage, robot hardware could become a negative-margin product, serving only as a carrier to put more terminals in the physical world, while value migrates almost entirely to the 'embodied brain'. Gao defined generalization as 'training cost': the higher the generalization, the lower the training cost. Xinghaitu now needs about 10 hours of post-training for a new long-horizon task, and aims to reduce that to one hour. The company has deployed fully autonomous robot front warehouses with JD.com.
At IJCAI 2026, Arizona State University assistant professor Hua Wei offered a cautionary academic perspective. He drew an analogy between large-model agents and reinforcement learning's 'sim-to-real' gap, noting that a model that works well in English can fail sharply in Chinese or low-resource languages. He proposed using domain randomization, a technique from robotics, by randomly perturbing prompts, action spaces and reward spaces during training. Experiments showed that a 3B-parameter model with such randomization could outperform a 32B model in cross-lingual tasks. Wei emphasized uncertainty quantification and human-in-the-loop as necessary safeguards, saying machines must know when they are uncertain and seek human help.