Shanghai Forum Debates World Models as AI's Route Into the Physical World
At the Inclusion·Bund Summit in Shanghai, researchers and robotics companies discussed whether world models can move AI from digital tasks into physical environments, presenting new models and data platforms while leaving reliability, data cost and commercialization unresolved.
Ma Yi, founding dean of HKU School of Computing and Data Science, opened the forum by describing the school's discipline building and its teaching and research layout in Shanghai. He said AI's rapid development is changing the relationship between talent training, scientific research and industrial translation, and that areas once separated now need to be closely connected in the same innovation process. Ma said HKU hopes to combine its international education and academic research strengths with Shanghai's industry, talent and innovation resources to explore a university-industry-research model suited to the AI era and to promote AI applications in healthcare, health, energy and materials. He encouraged university faculty to turn frontier research into entrepreneurial practice. On world models, Ma said perception and modeling of the three-dimensional and four-dimensional world have long been important problems in computer vision. Today's world models go beyond traditional 3D reconstruction and object recognition, pointing toward robots' perception, cognition, learning and interaction in open physical environments, which creates new space for university research and industry cooperation.
Wang Xiaogang, chairman of Daxiao Robotics, co-founder and executive director of SenseTime, and professor at the Chinese University of Hong Kong's Department of Electronic Engineering, spoke on Physical AI's Awakening Moment. He said embodied intelligence must overcome limits in both data and models to move from specific tasks to broader applications. Collecting real-machine data and training models repeatedly around a single task cannot economically cover changing real-world needs. Daxiao's ACE R&D paradigm uses wearable sensing devices to collect human behavior data in real production and living environments and fuses it with real-machine data, reducing dependence on large amounts of task-specific and body-specific data. For models, Wang introduced the Kaiwu world model's integrated architecture of understanding, generation and prediction. It judges environment and task states through understanding, simulates possible consequences of different actions through generation, and supports robot control through prediction, with the three abilities sharing features and forming a feedback loop. Using long-horizon tasks such as laundry as an example, he said a robot must not only recognize spatial relationships and decompose tasks but also judge whether each step is complete and adjust later actions. Data quality is a foundation. Wang said data from real scenes, including multi-view, force-tactile and interaction feedback, help models learn object properties and physical rules. Diverse behaviors and failure cases in open environments are more valuable for generalization than repeating the same task in a fixed environment. He also introduced environmental data collection devices, an automated data production platform, ACE-Data-0 open-source data and related evaluation work. In industrial applications, Daxiao is adapting different robot bodies through a one-brain-multiple-forms approach and advancing deployment in instant retail, hotel laundry and autonomous operations in open scenes. Wang said commercial deployment both tests model capability and provides real data for iteration.
Luo Yihang, co-founder and CEO of Shengshu Technology, gave a talk titled Seeing the Future Before Acting: From the First Principles of World Models. He argued that for AI to act in the physical world, it must first perceive the environment, infer the future and then continuously correct decisions based on action results. A general world model should have a complete closed loop of understanding, prediction and action; video generation, spatial reconstruction or policy control alone covers only part of that capability. Luo introduced Shengshu's five-level evolution route: simulating and generating the world, real-time interaction, driving physical action, developing into a world agent capable of autonomous planning and iteration, and finally exploring multi-agent autonomous collaboration. He linked this to practices including Vidu Q, Vidu S, Motus and Motubrain, describing the team's extension from digital content generation to embodied action and its exploration of cross-body adaptation, long-horizon task execution and dynamic decision-making. At the forum, Luo released Motus2, a general world model for dexterous manipulation. The model incorporates tactile information into unified modeling and combines action sampling, consequence prediction, value evaluation and policy improvement, so successful, suboptimal and failed actions can all become learning signals. It also introduces a memory mechanism to handle occlusion, environmental changes and state maintenance in continuous operations. Luo said dexterous manipulation is a severe test for world models entering the physical world, because after a robot touches an object it must judge force, adjust direction and anticipate consequences. He said current exploration remains insufficient, with long-term memory, online learning, joint evaluation and efficient deployment still challenges on the path to autonomous agents.
The roundtable was moderated by Luo Ping, associate dean for AI research and technology transfer and associate professor at HKU School of Computing and Data Science. Panelists included Yang Yanchao, co-founder and CTO of Yisheng Technology and assistant professor at HKU School of Computing and Data Science; Shi Ye, founder of Shunshi Technology, researcher and assistant professor at ShanghaiTech University and head of YesAI Lab; Zhu Zheng, co-founder and chief scientist of Jijia Shijie and director of the Beijing Key Laboratory of General World Models; and Li Tianyu, co-founder and CEO of YuanCe Future. They discussed definitions, research motivations, risks, suitable scenarios and the impact of large model development.
On what kind of system counts as a world model, Li Tianyu drew a distinction between broad and narrow definitions. Broadly, models encoding cognition of the physical world can be included; narrowly, for embodied applications, the focus should be on whether a policy or robot can interact with the model. Zhu Zheng favored the video generation route for its accumulated work in data, architecture and training methods, but said a humanoid robot's world model must also consider coordination among the brain, cerebellum and dexterous hands. Shi Ye said traditional physical simulation and data-driven world models are complementary: the former contains explicit physical laws, while the latter can absorb environmental changes and interaction experience from data. Yang Yanchao proposed three key conditions: understanding and predicting the future, extracting predictable structure from data and forming memory, and relying on memory for continuous learning and evolution.
Asked which real problems are pushing AI from generation toward understanding, Li Tianyu said physical AI has little tolerance for error, and real deployment demands reliable continuous operation, requiring deeper environmental understanding. Zhu Zheng said that although world models have drawn attention, they remain far from general usability, and scaling data and models must consider computing costs and commercial feasibility. Shi Ye used the example of a robot pouring water or tea: learning an action on fixed data does not mean understanding a new task, and if the goal or environment changes, a system relying on existing trajectories may continue executing the wrong action. Yang Yanchao compared the problem to a child learning a new toy: a model may generate realistic images but may not learn a new skill after watching one demonstration and making a few attempts as a child would. He said such rapid understanding and learning is especially important in changing home environments.
On where world models may be most overestimated, Yang Yanchao warned against equating accurate generation with correct understanding. He said realistic images do not necessarily mean better robot interaction, and whether a model has learned structures useful for action deserves more attention. Shi Ye said the meaning and technical paradigm of world models will keep evolving; adding more modalities may raise the capability ceiling but also brings extra burdens in computing, data and architecture design. Zhu Zheng called the commercial closed loop a key challenge: whether world models can find clear paying demand and use revenue and financing to support next-generation R&D will determine whether the route can continue. Li Tianyu focused on the cost of obtaining high-quality training information. He said representations such as video, 3D, depth and touch are not naturally low-cost, and scaling low-noise, physics-rich data is a long-term problem.
On which scenarios need world models most, Yang Yanchao said the key considerations are trial-and-error cost and data acquisition cost. When real-world trial and error is expensive and data are hard to obtain, predicting action consequences before execution is more valuable; for simple, low-cost tasks, it may not be necessary to call a world model for every decision. Shi Ye added that some industrial tasks are not only expensive for data collection but also costly to reset environments, and world models can help in these steps. They may also support feedback learning and continuous iteration after robot deployment. Zhu Zheng emphasized demand in open environments such as homes, where tasks and objects vary and models need to adapt quickly from few demonstrations. For industrial or retail tasks that are relatively fixed in the short term, VLA combined with pretraining and post-training already has some solving capability, so the necessity of world models should be judged case by case. Li Tianyu said visual world models are more likely to excel in simple, less cluttered environments, while reliability in complex backgrounds still needs improvement.
On the impact of large language model generalization, Yang Yanchao said progress in 3D understanding, coding and tool use brings new possibilities for physical intelligence, and the key is combining those abilities with memory and continual learning mechanisms. Shi Ye said stronger large models and agent tools could improve efficiency in data processing, R&D and automated workflows, but a gap remains between tool-calling ability and a robot directly understanding and executing physical tasks. Zhu Zheng said from an industrial competition perspective, foundation model companies have strong talent and resource reserves, and world model and embodied intelligence startups need to deliver technical and product value quickly. Li Tianyu said 3D asset generation and data cleaning are areas where large models can directly help. However, robot edge computing power must match the value it creates, and cloud large model technical paths cannot be simply copied to embodied systems.
In the forum summary, Luo Ping said world models may not be the only path to physical intelligence, but they have important potential for giving machines spatial imagination, common-sense reasoning and action prediction. On behalf of HKU School of Computing and Data Science, he invited industry to cooperate on scenario validation, joint R&D, technology transfer and ecosystem building to move research results into production lines and daily life. From predicting the world to reliably participating in it, the value of world models ultimately needs to be proved in real tasks. Whether they can lower adaptation costs for new scenarios, learn continuously from failures and create measurable benefits in long-term operation will be an important measure of the direction.
Editor's Summary The Shanghai forum gathered academics and companies to assess world models as a route from digital AI to physical action, with presentations on data collection, integrated understanding-generation-prediction architectures and a new dexterous manipulation model. Panelists agreed on the potential of world models for spatial imagination and action prediction but left definitions, technical routes and commercial viability unsettled. Reliability, data cost, long-term memory and deployment efficiency remain central challenges.