TUM Professor and New TrAct Paper Target AI's Deployment Gap
Leiphone reports: TUM's Ziyue Li and a new Li Fei-Fei/Wu Jiajun paper both tackle AI's gap between benchmark success and real-world deployment.
According to Leiphone, Li, who holds a PhD from HKUST and has worked at Bell Labs Stuttgart, Hong Kong MTR and SenseTime, used his IJCAI 2026 Early Career Spotlight talk to present what he calls the PDE triangle. The framework is not the partial differential equation, but Perception, Decision and Explanation. He argued that public-sector clients face incomplete, messy and uncontrollable data, while much academic research trains on clean data. He cited a McKinsey survey from about three years ago in which transportation and industry practitioners said true intelligence would take 10 to 20 years, and noted that no city in the world has long-term deployed reinforcement-learning-based traffic signal control.
On the perception side, Li said models must see all possible scenarios during training, including rain and accidents, because historical data often contains mostly sunny conditions. He described a partner city of more than 30 million people where only 120 intersections had sensors, and said his group used masked-node pretraining and virtual nodes to infer the missing data. On the decision side, Li listed several reinforcement-learning signal-control projects aimed at cross-city transfer, multi-agent cooperation, fast adaptation to new scenarios and engineer-friendly guidance, but he added that interpretability is the first of three obstacles that still stop deployment.
The arXiv paper, titled TrAct: Bridging Robot Control and Visual Prediction with Visual Tracks, also comes from the intersection of prediction and control. Leiphone said the authors include Li Fei-Fei and Jiajun Wu, who have worked together before; a month earlier they published a paper with ControlNet author Lvmin Zhang and others that rendered robot actions as pixel trajectories for a video world model.
The TrAct paper argues that robot controllers output low-dimensional action dialects - end-effector poses, joint angles, gripper commands - while world models need high-dimensional pixels. Directly conditioning a world model on actions, labeled AWM, forces the model to infer an entire physical-causal chain and can cause wrong gripper poses, disappearing objects and inconsistent camera views. TrAct instead tracks two sets of image keypoints: seven gripper points and 25 scene-grid points. These visual tracks are embodiment-agnostic: moving an object into a bowl looks like the same track whether the actor is a Franka, a UR5 or a human hand.
Leiphone described the architecture as a three-part pipeline. VLAT, built on the pre-trained flow-matching VLA π0.5, outputs paired action blocks and track blocks from the same policy sample; TWM, based on Stable Video Diffusion with a Temporal ControlNet branch, renders the tracks into predicted videos; and VLAC, a vision-language reward model built on InternVL2, scores each video by task completion. At inference, VLAT proposes 20 candidate pairs in simulation (16 on a real robot), TWM generates previews, VLAC picks the best one, and the robot executes the action paired with the winning track.
On the data side, TrAct mixes 76,000 DROID robot teleoperation episodes with about 150,000 EgoDex human first-person videos, a roughly 1:2 ratio. Human videos have no robot action labels, but they can supply visual tracks, so they participate in pretraining without being translated into actions. In the standard LIBERO benchmark, TrAct reached 98.3 percent versus π0.5's 96.8 percent. On the team's self-built LIBERO-INTEGRAL benchmark, which includes robustness tasks and cross-embodiment tests with a UR5 replacing the Franka Panda, TrAct scored 55 percent on average, more than double π0.5's 27 percent and six points above VLAT+AWM at 49 percent. The paper reports that TrAct led in all six task categories and was best or tied in 18 of 20 tasks, with non-overlapping confidence intervals against VLAT+AWM across three seeds.