AI News Feed
Market watch
Computer Vision

Wu Proposes Physics as the Common Ground for Vision, Sound and Touch at ECCV 2026

Stanford's Jiajun Wu told ECCV 2026 that vision, hearing and touch are projections of one set of physical properties onto different channels, and that fusing them under scarce data calls for physical principles rather than bigger architectures. He described distilling stiffness from video diffusion models and turning a single image into an interactive 3D scene.

According to Leiphone, the talk was delivered in the "Embodied Multimodal Reasoning in Physical Environments" special session at ECCV 2026 and was Wu's third major address in two months. On Aug. 20 he received the IJCAI 2026 Computers and Thought Award in Bremen, Germany, where he argued for encoding geometric and physical priors as "physical code" and stitching them into foundation models. On Sept. 8, at the conference's Robotic Manipulation Research workshop, he turned to manipulation and planning, describing how foundation models can discover atomic skills together with their preconditions and postconditions. Both of those lines ran from vision downstream to planning and acting; the Sept. 9 talk flattened the sensing side instead, placing vision, audio and touch on the same footing.

The problem Wu set out is how to fuse those modalities when data are scarce — entering a new domain, or facing objects never seen before, leaves almost no recordings of how they sound or how they feel. "Of course, you can just use a Transformer," he said, noting that this is the prevailing approach and that it works reasonably well. "But sometimes we wonder whether we can do a bit better, especially when we have very little data." His illustration: an object made of porcelain or plastic looks like plastic, feels like plastic, and sounds like plastic when struck.

PhysDreamer was the earliest of the works he described, an ECCV 2024 oral presentation that distills physical properties from video diffusion models. The task was action-conditioned prediction of how a static 3D object, such as a flower, would move once its physical parameters were inferred. The critical quantity was Young's modulus, the measure of stiffness, and Wu said the estimation is harder than it appears because an object is not homogeneous. "Even for a single flower, different parts have different physical properties," he said. Heuristic assignments of physical properties produce simulated behavior that is visibly unrealistic and highly variable, and no ground-truth annotations exist for a stiffness field.

The method represents a scene with 3D Gaussians, renders images, and sends them through a video generation model to obtain a reference video, while a candidate Young's modulus field is carried alongside an appearance field in a differentiable material point method pipeline. The simulated trajectory is rendered and compared with the reference, and because the whole chain is differentiable, the resulting loss updates the physical estimate. "Even if you only compute the loss in pixel space, it gives you a lot of useful information," Wu said. Later work replaced that pixel-space loss with a score-distillation-style objective, which he described as less brittle. Compared against motion captured from real flowers, the estimates produced more realistic motion than baselines, he said.

A subsequent line of work, WonderPlay, extends modeling beyond deformable objects to rigid bodies, fluids, granular materials and articulated objects, reconstructing an interactive 3D scene from a single input image. Forces and wind fields can be defined in 3D space; demonstrations in the talk showed a stick dragging through honey and cake, and wind from the left moving a hat, hair and smoke. Wu argued that the conventional route of state estimation followed by physical simulation is insufficient, because full state estimation is extremely difficult and interactions such as fluid against rigid body require heavy approximation. Video models supply the visual prior that simulators lack, he said, but both are constrained by the data available: generative systems train mostly on video from drones or games, so the actions they accept are largely navigation commands.

The same reasoning, Wu said, extends to sound and touch, through work on differentiable rendering of collision sounds and on flexible electronic skin for tactile sensing. His talk was titled "Physics-Grounded Multimodal Perception: From Vision and Sound to Touch."