World Labs unveils Atlas, touted as first multimodal world model
World Labs, co-founded by Fei-Fei Li, released Atlas, which it calls the first multimodal world model, creating 3D worlds from single images and enabling robot training simulations.
In a blog post, World Labs described Atlas as an “omni model” that can model the world, move the camera, and simulate space and time. From a single 2D image, it can produce up to one minute of 1440p video while keeping geometric consistency, and it can be viewed from any angle. The model can also take one to dozens of input images to reconstruct a real-world scene, generating new views or explicit 3D representations such as point clouds and Gaussian splats. In addition, Atlas can create images and 360-degree panoramas from text prompts, supporting complex instructions and a wide range of visual styles.
The underlying architecture is a multimodal autoregressive diffusion transformer, according to the company. Text, image, video and depth data are anchored in three-dimensional space to form a spatial context, and the model generates outputs from that context. World Labs said the design draws on techniques from large language models and modern video generation models, allowing it to use methods like KV cache and diffusion distillation.
Atlas is particularly aimed at robotics training. It can generate realistic RGB and depth data from a few photos of a physical space, so robots can practice in many simulated environments. This real-to-sim workflow was highlighted by Jim Fan, head of robotics at Nvidia, who said in a post quoted by Chinese tech outlet Quantum Bit that Atlas is a significant step for real-to-sim in robotics. Fei-Fei Li herself called the release a milestone for World Labs.
World Labs was founded in February 2024 by Li, who has argued that spatial intelligence is necessary for artificial general intelligence. According to SiliconANGLE, the startup has raised $1.2 billion from investors including Nvidia, AMD and Autodesk. The company said it evaluated Atlas on benchmarks for camera-control generation and 3D reconstruction. It outperformed state-of-the-art video models in camera-path adherence, with a bigger advantage on complex trajectories, and exceeded the best open-source 3D reconstruction models in sparse-input reconstruction. Atlas is now in early access with some partners, and other researchers can apply through the company's website.
SiliconANGLE noted that Atlas enters a crowded world-model market, with startups such as Odyssey, AMI Labs and Niantic Spatial developing similar technologies. World Labs said Atlas will serve as the foundation for future versions of its Marble product and other offerings.