AI News Feed
Market watch
Robotics

WorldArena 2.0 Challenge at IROS 2026 Puts World Models to Real-Robot Test

WorldArena 2.0 Challenge at IROS 2026 unveiled final rankings on Sept. 16. Three tracks tested video quality, online RL environments and real-robot manipulation. Visincept, Zhongguancun Academy and Xiaomi Robotics led their tracks.

WorldArena 1.0 evaluated 14 representative models on two broad capabilities: the visual quality of predicted future video and whether those predictions were useful for data generation, policy evaluation, and action planning. Its 15 video metrics covered visual quality, motion quality, content consistency, physics adherence, 3D accuracy, and controllability. But the earlier benchmark remained mainly visual, offline, and simulation-based. It did not systematically test tactile feedback, online interaction, or execution on real robots, where noise, delay, and error accumulation matter.

WorldArena 2.0 added three dimensions: modality, functionality, and platform. It expanded from vision-only inputs to visual-tactile perception; from offline evaluation to online reinforcement learning; and from simulation environments such as RoboTwin 2.0 and LIBERO to real-robot testing. The Challenge used three tracks: Track 1 for video quality, Track 2 for world models as reinforcement-learning environments, and Track 3 for real-robot operation. Tracks 2 and 3 were new.

Track 1 had 70 official tasks: 50 in-distribution clean scenes on white tabletops and 20 out-of-distribution scenes with colored or textured tabletops. All models were tested on the same dataset. The track assessed long-horizon and OOD video generation, with 15 metrics spanning image quality, motion, consistency, interaction accuracy, 3D geometry, and instruction following. JEPA Similarity was also included, using V-JEPA to compare feature distributions between generated and reference videos. Eighty-one models received final scores. Visincept's WorldIncept, which explicitly follows the JEPA route, ranked first with 72.33 points; the Tongji University spatial intelligence team's BWM-Turbo was second with 71.49, a gap of 0.84 points. WorldIncept led in physics adherence, 3D accuracy, and semantic alignment, where it scored 90.11. Visincept was incubated by the IDEA Research Institute and has released DINO-X visual models, DINO-XGrasp, and an EgoTwin data engine with Baidu AI Cloud, building robot dynamics modeling on hundred-thousand-hour ego and teleoperation data. BWM-Turbo led in background consistency and JEPA Similarity, emphasizing stable long-horizon generation.

Track 2 tested whether a world model can serve as an online reinforcement-learning environment. Policies generated actions, the world model predicted future states, rewards were computed from those predictions, and the updated policy was then tested in RoboTwin 2.0 on the Adjust Bottle task, which requires fine pose control and stable grasping. Among 47 evaluated models, Beijing Zhongguancun Academy's MW2 ranked first with 72.93 points, and Chenhunxian Technology's TTWM was second with 72.40, a gap of 0.53 points. They were the only two entries above 72. TTWM is the competition version of Chenhunxian's goal-causal world model GCWM, which models how actions change future states and uses optimizations around target regions, spatial geometry, and temporal consistency. Chenhunxian also follows the JEPA route. A baseline VLA π0.5 achieved a 55.46 percent success rate; after policy optimization with GCWM as the reinforcement-learning environment, the VLA model's success rate rose to 72.40 percent.

Track 3 moved evaluation to real robots, covering AgileX dual-arm and Franka Panda single-arm platforms. Track 3.1 tested visual-tactile manipulation on AgileX with three tasks: picking up potato chips, peeling a cucumber, and inserting a two-prong plug. Track 3.2 tested vision-driven manipulation. On AgileX, it included wiping a table, pouring water, clearing a table, clearing a table by instruction, folding clothes, and folding a paper box; Franka also covered wiping, pouring, and clearing. Tasks were scored by real-robot success rate, and the overall leaderboard combined task difficulty and weighted visual and visual-tactile results. Robots had to complete tasks within a set number of steps without triggering safety interventions.

In the overall Track 3 leaderboard, Xiaomi Robotics MiRobot team's ViTacX ranked first with 76.64 points, while ShanghaiTech University's OmniFlow team was second with 67.37. ViTacX's main advantage came from the vision-only sub-track, where it scored 85.00 and ranked first, taking the highest single-task scores in wiping a table, pouring water, folding clothes, and folding a paper box; the first three were 100 points each. In Track 3.1 visual-tactile sub-track, ViTacX ranked third with 66.67. OmniFlow stayed second in both the vision-only and visual-tactile sub-leaderboards. Xiaomi Robotics has focused on embodied foundation models, VLA, and real-robot execution, while ShanghaiTech's automation and robotics center has worked on tactile sensing, teleoperation, imitation learning, reinforcement learning, and multi-contact motion.

According to Leiphone, the three tracks were ranked and awarded independently, with total prize money of 7,000 dollars based on official award amounts. The final results also showed a split between academic and industry teams: in each of the three tracks, the top two included one academic and one industry entry. The Tsinghua organizers told AI Tech Review that the benchmark's move toward online interaction and real-robot execution required new design choices and collaboration, and that the evaluation still has boundaries, according to Leiphone.