PHYBOT Humanoid Completes 40-Plus Shot Badminton Rally at IROS 2026
At IROS 2026, China's PHYBOT presented a humanoid badminton system that separates tactical decisions from whole-body execution, with a 1.35-meter robot trading more than 40 shots in a rally with a human.
Leiphone reported that badminton is unusually hard for robots. The shuttlecock can leave the racket at several hundred kilometers per hour but decelerates sharply because of air drag; its path is asymmetric, and its cork-and-feather construction makes it flip in flight. The sport therefore requires a division between a cerebellum for whole-body control and a cerebrum for strategy. Whole-body control has advanced from stable walking to humanlike running, jumping and dancing, making racket sports a new test. Previous work has favored layered architectures, with high-level planners and low-level trackers, and human priors from motion capture, video and amateur data. Most public work addresses interception; strategy remains less explored.
In November 2025, Liu was first author of a paper that trained a unified whole-body controller with a three-stage curriculum, without motion priors or expert demonstrations. The real robot reached a maximum shot speed of 19.1 meters per second. The team later added human reference motions and multiple discriminators for topspin, backspin, forehand and backhand styles, plus a learned skill selector that chooses the best-performing skill for each incoming shuttle.
At the higher level, a GRU policy decides the hitting technique and target placement. The team used league-style zero-sum multi-agent reinforcement learning to create stronger opponents in simulation and expose policy weaknesses. A fixed lower-level controller executes the chosen technique and placement. Leiphone reported that the same framework transferred to table tennis and soccer in about one week and was also taken to a shopping mall pop-up venue.
In his talk, titled Humanoid Badminton: Bringing Athletic Robots from the Lab to the Court, Liu described racket sports as different from soccer or basketball, where contact is brief. Racket sports have a much smaller target, precision demands roughly an order of magnitude higher, and require repeated reliable hitting. Interaction through a racket amplifies small joint errors, and high ball speeds leave a narrow striking window. Badminton adds strong air drag, rapid nonlinear deceleration and frequent acceleration and deceleration. Liu also cited China's large badminton-playing population.
The talk outlined deployment gaps. Robots must handle noise, latency, disturbances and out-of-distribution inputs. A deployable system also needs navigation, shuttle retrieval, fall recovery and long-term stability. The best solution is often not theoretically optimal but reliable in the real world.
Early work used an annealing-style curriculum for whole-body control and achieved a first field demonstration of a humanoid playing badminton in a motion-capture environment. Training began with a lower-limb movement reward, which was gradually removed as the hitting reward became dominant, preserving learned movement skills. The first laboratory hitting test was completed in November 2025. The system was later extended to multiple hitting skills and target placement. A return reward was introduced after the robot learned to hit, and a one-hot encoding of the target area was added to the policy observation. The learned skill selector, a Q-value function, predicted task utility, defined as the product of hit reward and return-placement reward, and executed the highest-utility skill. In simulation, the system without the selector ended rallies early, while the version with it continued exchanges.
The team said its system achieved the longest rally in legged-robot badminton. Video showed Liu playing against the robot. The 1.35-meter robot covered a 4.4-by-4.4-meter court, and the selector chose techniques such as backhand when optimal. Liu said the robot could not be shipped to the United States for a live demonstration.
For strategy, the high-level problem was formulated as a zero-sum multi-agent reinforcement learning game. A main agent trained against historical opponents, while prioritized fictitious self-play focused on opponents the policy struggled to beat. When the win rate was high enough, a snapshot was saved and training continued. The high-level policy observed robot, hitting and opponent states and previous decisions, with a GRU outputting two discrete decisions: hitting technique and return placement. League training used a main agent, a main exploiter to find weaknesses and a league exploiter to introduce diverse strategies, intended to avoid cyclic dominance like rock-paper-scissors. Training charts showed improvement against strong historical opponents while new weaknesses emerged and training focused on difficult ones.
The high-level results remain limited to simulation. The next step is validation on real robots and understanding how advanced strategies evolve. In a video, agent 12 outperformed agent 10, an earlier version. The team also tested transfer beyond badminton.