Chinese AI Startup Shows Similar World Model Six Months Ahead of Atlas, Tops Benchmark
World Labs released Atlas on Sept. 2, but Chinese startup Yingsu had already open-sourced a comparable world model in March and later ranked first in WorldArena 2.0. Yingsu also launched an open data initiative.
Atlas can integrate text, images, video and 3D information into a unified spatial context, QbitAI said. With one or a few images, it reconstructs a 3D scene, generates new viewpoints and moves a camera along specified trajectories while preserving spatial consistency. Fei-Fei Li described the release as a milestone.
Yingsu's open-source model, InSpatio-World, works from an ordinary video and builds an explorable dynamic 4D scene from limited spatiotemporal observations. Like Atlas, it lets users alter the original camera path and look at areas not captured before; unlike Atlas, it explicitly models time, so users can shift the observation moment within existing events and re-watch the motion from new viewpoints. Yingsu said an upgraded version would be released soon and remain open-source. Both models point to world models generating observations of an ongoing spatial or spatiotemporal state from different positions and views, rather than simply rendering plausible video frames.
On Aug. 27, Yingsu's InSpatio-Curious topped the first phase of the WorldArena 2.0 leaderboard with 66.11 points out of 77 models. It was first in Track 1 and in Physics Adherence, Trajectory Accuracy, JEPA Similarity and Depth Accuracy. Its trajectory accuracy was 64.89, 3.02 points over the runner-up; depth accuracy was 99.29; JEPA similarity was 98.36; physics adherence was 73.24. The model's image quality score was only 60.64, nearly 9 points below the overall runner-up, indicating the win came from deeper capabilities rather than visual polish.
WorldArena was launched by Tsinghua University, Shanghai Jiao Tong University, the University of Hong Kong and others to benchmark embodied world models, according to QbitAI. Version 2.0 checks whether movement trajectories are accurate, spatial relationships remain stable, and objects continue according to physical rules after interaction, rather than only judging picture clarity. The report also noted that the benchmark currently centers on standardized simulation and robot manipulation tasks, and Track 1 scores rely on 15 generative and visual-agent indicators without directly measuring real-world forces, friction, mass or inertia. A high ranking under a fixed test set does not yet amount to qualification in general spatial intelligence.
On Aug. 29, at the Second China Spatial Intelligence Conference in Wuhan, Yingsu and more than 20 university teams and industry organizations launched SIDO, a large-scale open data plan for spatial intelligence. SIDO aims to connect data producers, model trainers and users across LiDAR, robotics, 3D generation, digital space and industrial applications. Over two years, it plans to build a data foundation containing millions of 3D/4D scenes and to develop benchmarks around 3D reconstruction and generation, scene understanding, spatial reasoning, dynamic prediction and embodied planning.