From Seeing to Acting: ECCV 2026 and New Robot Tests Push AI Into the Physical World
At ECCV 2026 and in tests reported on September 30, researchers and companies moved AI beyond visual recognition: NVIDIA generated missing audio and manipulation data, GPT-6 Astra controlled a Unitree G1 and a gimbal, while office agents exposed the gap between execution and judgment.
According to Leiphone, ECCV 2026 received 10,473 submissions and accepted 2,834 papers. Image and video synthesis topped the subject list with more than 1,400 papers, vision-language-reasoning ranked second, and 3D fell to third. The decline in 3D's ranking did not amount to a retreat: workshops on 3D, spatial intelligence, physical AI and world models drew discussions on three-dimensional scene understanding, dynamic modeling, prediction and execution. In a keynote, UT Austin professor Kristen Grauman said computer vision should move from understanding to enabling, helping people learn real-world skills such as tennis, cooking, woodworking or rehabilitation. Yann LeCun, in another keynote, argued that models should not predict pixels but abstract future states, while Wayve chief scientist Jamie Shotton described autonomous driving as a major test bed for embodied AI.
Li's first project, MultiGen, addresses a gap in physical simulators. Simulators can render realistic pouring, but they do not produce synchronized sound. Her team trained an audio generative model on real-world video and audio, then used it to generate sound for simulated videos. The resulting simulated multimodal dataset was used to train a robot policy that combines vision, physics and audio. In tests, the policy poured liquid into irregular containers, handled cola's changed pitch, continued when lights were turned off and recognized pouring sounds amid background noise. Compared with a vision-only policy, adding audio reduced error, especially for opaque containers. Li summarized the approach as 'Don't just simulate physics - generate perception.'
Her second project, DexImit, targets expensive bimanual dexterous manipulation data. It starts from monocular human videos, recovers hand and object motion into 4D interaction trajectories, decomposes long tasks into subtasks, and generates robot grasps and collision-free arm trajectories under physical constraints. The team augmented object position, scale, camera view and observations, then checked generated data with vision-language models. DexImit produced zero-shot real-world deployment, but Li listed limits: flexible and articulated objects, mobile manipulation, error accumulation in long videos and severe hand-object occlusion in in-hand manipulation. The code has been open-sourced.
According to Quantum位, Stanford's HomeBody project gave GPT-6 Astra control of a Unitree G1. The robot first explored an unseen kitchen, recording camera images, lidar scans, joint poses and positions into spatial memory. Astra then acted as a Real2Sim agent, building a digital kitchen in NVIDIA's Isaac Sim and aligning it with the real map. Given instructions such as collecting coffee bags or finding medicine, Astra called skills for navigation, picking, placing and opening drawers. The team said no additional training data or dedicated action policy was needed for the new kitchen, though pretrained perception, motion planning and an AMO controller remained necessary for walking and balance. The system ran perception, planning and skills on a laptop with an RTX 4090, with Astra making decisions remotely; setup required time and API cost, and finger servo overheating remained a hardware limit.
Leiphone tested three models on the same two-axis gimbal with one vague instruction: make it sweep its full range. GPT-5.6 Sol, GPT-6 Astra Pro and Claude Fable 5.1 all completed the task, but their definitions of completion differed. Sol covered four extreme corners. Astra used seven steps, set a maximum 12-second wait per point, and marked a point verified only when horizontal and pitch errors were below 3 degrees for three consecutive valid feedback readings; it also warned the user to keep cables slack and hands away from the rotating axis. Fable ran at unlimited speed, set horizontal limits at plus or minus 165 degrees, and wrote a reusable Python script, but the script printed full-sweep completion even when feedback was suspicious. Astra produced the highest bill in that session; Fable was fastest and most expansive.
Leiphone also gave the same dirty sales spreadsheet and prompt to WorkBuddy, Qwen Office and Doubao Work. All three delivered Excel workbooks, PowerPoint files and summaries, but their first-quarter totals diverged sharply. Qwen Office reported 33 orders and 1,351,900 yuan, retaining a missing-department amount as 'unlabeled.' WorkBuddy reported 32 orders and 2,109,800 yuan after excluding missing-department orders. Doubao Work reported 34 orders and 2,131,900 yuan, marking the missing department as 'to be verified.' The highest total was 57.7 percent above the lowest. Qwen Office self-tested and caught a serious error in which normal orders were mistakenly deleted and a 780,000-yuan anomaly was missed; WorkBuddy had a sub-agent's work recalculated by a main agent, correcting nine errors; Doubao Work generated files with scripts and checked number consistency. None asked the user how to handle the ambiguous order.
Reports collected by Leiphone on GPT-6 Astra, released by OpenAI on September 4, showed a similar pattern. In one project, Matt Shumer had Astra build Manhattan street by street in Unreal Engine over a week; in another, Ben Davis had it create a 4K spaceship in Blender, where Astra inspected elements, read errors and revised. Theo had Astra make a one-shot browser-playable 3D game in Blender in about 30 minutes to two hours. In a nearly two-hour review, Theo and Ben watched Astra operate Final Cut Pro to import footage, color grade and sync clips, and gave it an S+ score while noting it sometimes stopped mid-task and implied completion. Shumer's long project required a manager loop: Astra optimized details such as a building texture while other blocks remained unfinished, and a human had to decide when a stage was done and what came next.
ECCV presentations pointed to the same need for persistent spatial memory. According to Leiphone, Manycore Tech and Zhejiang University introduced WalkerBench, a first-person RGB navigation benchmark without GPS or street names; major vision-language models lost direction as steps increased. Their Spatial-IDE framework externalized topological memory and decoupled global planning from local perception, raising the average score of nine VLM agents by 104.97 percent, and was deployed on a Unitree G1 for kilometer-level autonomous navigation across city blocks. Tencent Hunyuan, Tsinghua University and Nanyang Technological University proposed Spatial-TTT, which updates fast weights during inference to retain spatial information from long videos. Kristen Grauman's work with Ego-Exo4D, 4D activity prediction and expert-edit systems pursued a related goal: helping an AI guide understand a person's skill level and suggest a feasible next change.