AI News Feed
Market watch
Large Language Models

HiDream.ai Releases Embodied World Model HiDream-O1-Embodied, Tops Robustness Benchmark

HiDream.ai unveiled HiDream-O1-Embodied, an embodied world model that ranked first in the Robustness sub-leaderboard of RoboColiseum with a score of 0.692, completing its native all-modal technology loop.

HiDream-O1-Embodied was evaluated for the first time on RoboColiseum, an embodied-intelligence benchmark platform. It topped the Robustness sub-leaderboard with an average score of 0.692. The platform assesses models through 78 high-fidelity simulated tasks across four dimensions: instruction following, spatial understanding, robustness and general manipulation. Robustness is widely seen as the most demanding dimension because it changes backgrounds, lighting, materials, robot start states, camera positions and image quality, while also rewriting instructions in varied ways.

The model’s performance draws on what the company calls three breakthroughs. Its language understanding goes beyond keyword matching and recognizes equivalent instructions regardless of phrasing or sentence structure. Its visual perception processes multiple camera views in parallel, so when one view is misaligned or blocked, other views still let the model understand the scene and continue the task. Its high-tolerance mechanism exposes the model to deliberately imperfect conditions during training, including changing light, corrupted images, occlusions and signal fluctuations, enabling it to act on incomplete clues.

HiDream.ai CTO Yao Ting said the company believes a complete world model foundation must combine native all-modal expression, causal reasoning and physical world construction. “The release of HiDream-O1-Embodied is a key milestone in this technical strategy moving from the simulated world to the real world,” he said.

The company also credited a data production model it calls a “real foundation plus generative augmentation” approach. In a collaboration with Noitom, real high-precision human motion-capture data serve as the base, and HiDream.ai’s native all-modal capability expands each real clip hundreds of times by generating variant videos that alter background, lighting and object shapes without breaking physical constraints. The model thereby acts as both test taker and question setter, producing training samples that address its own weaknesses and create a self-reinforcing loop.

The announcement comes less than a month after HiDream.ai released HiDream-O1-World, an interactive world model that topped the Navi sub-list of the WBench benchmark with an average score of 80.9. The company said the interactive model handles understanding and reasoning in the digital world, while the embodied model handles operation and execution in the physical world.

Founder and CEO Mei Tao was quoted as saying the next generation of AI competition will move from single-modality to multi-modality and eventually to native all-modal models. With HiDream-O1-Image, HiDream-O1-World and HiDream-O1-Embodied, the company is forming a model family that covers visual, interactive and embodied capabilities.