Xingdong Jiyuan's VPP2 Tops RoboDojo Benchmark, Beats GPT-6-Astra and NVIDIA GR00T
QbitAI reports that Chinese robotics company Xingdong Jiyuan's World Action Model VPP2 has ranked first on the RoboDojo simulation benchmark, beating GPT-6-Astra, π0.5 and NVIDIA GR00T-N1.7. VPP2 also led real-robot zero-shot tasks and several generalization tests.
RoboDojo, led by the University of Hong Kong's MMLab and built with nearly 20 top academic institutions, includes 42 dual-arm manipulation tasks across five dimensions: generalization, precise manipulation, long-horizon tasks, memory and open-vocabulary instruction understanding. VPP2 placed first overall and also first in generalization, precise manipulation and memory, according to the report.
The report said VPP2 achieved the result without additional data or enhancement methods such as Agent RSI, relying only on standard datasets. In the RoboDojo simulation benchmark's post-training evaluation, GPT-6-Astra had an average success rate of 22.48 percent and an average score of 28.97, leaving VPP2 ahead by 9.78 percentage points and 10.29 points. Success rate measures whether the robot completes a task, while average score reflects how far it progresses even when it does not finish. Other cited competitors included Physical Intelligence's π0.5 and NVIDIA's GR00T-N1.7.
Xingdong Jiyuan also deployed VPP2 on a real ALOHA dual-arm robot for 10 types of zero-shot manipulation tasks, including grasping, placing, stacking, folding and pouring. VPP2 reached an average success rate of 58.5 percent, higher than π0.5's 40 percent, and recorded the best result in nine of the 10 task categories, the report said.
VPP2 follows the World Action Model route, which uses video prediction to understand how the physical world will change and then converts that prediction into robot actions. QbitAI reported that Xingdong Jiyuan addressed a longstanding problem in this route: good video prediction does not necessarily mean a robot can perform tasks well. The company decoupled video prediction and action learning and trained them in stages rather than jointly.
The first stage is event-level continued video pretraining, in which the model learns to predict a complete operation process. The second is fixed-duration video post-training and distillation, intended to make prediction fast enough for real-time robot execution. The third is action expert training, which turns video prediction into concrete actions while trying to preserve generalization.
The model is based on Alibaba's open-source Wan2.1-I2V-14B and integrates robot manipulation, human activity and general video data covering different robot bodies and manipulation methods. The team segmented operation processes into semantically complete clips and paired them with detailed descriptions, specifying which arm, cup or box is involved rather than using vague commands such as put the cup in the box. In an instruction-following test on robot manipulation videos, the 14-billion-parameter VPP2 reached 90 percent success, while the 64-billion-parameter Cosmos3 model reached 78 percent, according to the report.
To improve speed, the team adjusted the event-level prediction model to predict fixed eight-second video segments and used consistency distillation to compress multi-step computation into single-step generation. Predicting visual changes for the next eight seconds takes about 0.12 seconds. For action generation, VPP2 uses a 0.9-billion-parameter diffusion Transformer, Action DiT, under an MoT architecture. During early action training, the Video DiT's base parameters were frozen and adapted only through LoRA to avoid damaging prediction generalization. Video prediction takes about 0.12 seconds and the action expert about 0.1 seconds, for an overall action segment generation latency of about 0.22 seconds, the report said.
On LIBERO-Pro, which tests operation after changes in object positions and task requirements, VPP2 achieved an overall success rate of 45.0 percent, while the highest baseline was 11.0 percent. On LIBERO-OOD, which tests compositional generalization by recombining familiar objects, layouts and task goals into unseen tasks, VPP2 reached 63.9 percent.
QbitAI reported that a higher-level planner can work with VPP2. In the article's description, GPT handles understanding, reasoning and planning, while VPP2 predicts physical changes and turns plans into actions. The team used a VLM high-level planner for semantic understanding, memory and task decomposition, then handed subtasks to VPP2. In RoboDojo long-horizon tests, adding VLM subtask planning raised average success from 27.6 percent to 57.6 percent.
The report said Xingdong Jiyuan is pursuing a full-stack strategy covering the brain, the robot body and dexterous hands, aiming to control not only training but also deployment and feedback.