IROS 2026 Closes in Pittsburgh With Long-Term Memory Best Paper as 1,933 Papers Show Robotics Absorbing AI
IROS 2026 in Pittsburgh awarded its best paper to LT-Mem, a long-term memory system, while humanoid tennis, tray transport and a Stanford talk on structured representations showed robotics integrating foundation models with geometry, control and memory.
The Best Paper Award went to Yumin Lee, Hyoseok Ju and Giseop Kim for LT-Mem: Volatility-Aware Spatio-Temporal Memory for Lifelong Scene Understanding. The paper addresses temporal amnesia, where a robot updates its map and loses older states. LT-Mem uses layered Live Memory, Delta Memory and Meta Memory, and a volatility measure to track which objects change often and which remain stable over long periods.
The Best Student Paper Award went to Pei-An Hsieh, Fengjun Yang, Nikolai Matni and M. Ani Hsieh for Flatness-Preserving Residual Learning for Real-Time Tight Quadrotor Formation Flight. The work uses physics-informed residual dynamics learning to compensate for downwash and aerodynamic interference when quadrotors fly very close, while preserving differential flatness for trajectory planning and control.
Two humanoid awards drew attention. Zhang Zhikai and co-authors won the Best Entertainment and Amusement Paper Award for Learning Athletic Humanoid Tennis Skills from Imperfect Human Motion Data, which lets a humanoid learn tennis skills from imperfect human motion data. Anlun Huang and co-authors won the Best Paper Award on Mobile Manipulation for SteadyTray, which separates locomotion from payload stabilization and uses a residual policy to keep a tray stable while a humanoid walks.
The tennis work was also detailed in a separate report on LATENT, a project from Galbot, Tsinghua University, Peking University, Shanghai Qi Zhi Institute and Shanghai AI Laboratory. The team collected about five hours of motion data from five amateur tennis players in a 3-by-5-meter capture area, focusing on forehands, backhands and footwork rather than complete matches. It trained a motion tracker for a 29-degree-of-freedom Unitree G1, compressed human motion into a latent action space, and added a Latent Action Barrier to keep reinforcement learning near learned movement patterns. In simulation, the team randomized robot and ball dynamics, observation noise, frame loss and delay. Real-robot tests relied on motion capture, with more than 50 cameras covering a 19-by-15-meter area.
Also on Sept. 27, Stanford assistant professor Wu Jiajun gave an invited talk, Building Physical Agents via Structured Representations, at the RoBoWoMo workshop. Using a kitchen scene in which a student leaves a soapy dish for a robot, Wu described how small state differences can change a task: a clean dish should be put away, a dish without soap may need soap before washing, and a dish hiding dirt may require wiping the table. He argued that structured representations, learned predicates, atomic actions with preconditions and effects, and neural policies can serve as scaffolding for robot reasoning and action, not as a fixed final answer.
A review of the 1,933 accepted papers found that robot learning and embodied AI appeared in about 809 papers, navigation and planning in 564, perception and vision in 556, control and dynamics in 546, manipulation in 520, and humanoid and legged robotics in about 213. VLA, VLM, LLM and foundation-model papers numbered about 162, including about 82 on VLA. Reasoning and memory appeared in about 119 papers. The distribution, according to the review, points to foundation models being embedded in planning, perception, control and manipulation rather than collapsing robotics into one AI problem.
Specific papers reflected that pattern. A study on LLMs for task and motion planning with PDDLStream found that geometric constraints and motion feasibility still require traditional planners. GeoVLA added depth and 3D geometry to a vision-language-action model. Training Fast Robot Policies with Slow Foundation Models used large models in training and a lighter policy for real-time control. Shallow-π distilled a flow-based VLA for faster inference. Other work addressed camera adaptation, long-term memory for VLA agents, modular monitoring and replanning, visuo-tactile grasping, and 3D Gaussian scene memory. Special awards covered vision-language navigation, a redesigned robotic zipper for self-donning clothing, agricultural 3D reconstruction, and contact-aided underwater localization.