Stanford and Caltech Connect GPT Astra to Unitree G1 in Layered HomeBody Project
Stanford's Movement Lab and Caltech have connected GPT Astra to a Unitree G1 in the HomeBody project, letting the humanoid explore a kitchen, remember object locations and execute long-horizon tasks through local skills rather than end-to-end action generation.
The project diverges from end-to-end vision-language-action models. HomeBody builds an explicit interface between the model and the robot body. Astra issues structured tool calls that select a skill and provide the parameters that skill needs. Pick accepts a normalized two-dimensional image point from 0 to 1000 and specifies left or right hand. Navigate accepts a two-dimensional map target point and orientation. Place uses a three-dimensional release position in the robot torso frame, an execution hand and a release distance. Drawer operation packages visual alignment, hooking the handle and walking backward into a single skill.
The design deliberately prevents the VLM from outputting continuous end-effector trajectories. Astra expresses which target to manipulate in semantic space, and the skill interface converts that intent into geometric constraints. Astra does not need to know every degree of freedom of the G1 or how the inverse-kinematics solver works. It only needs to know which capabilities are available and what parameters each capability requires. Grasping can be handled analytically, navigation can use a conventional planner, and a new skill can come from reinforcement learning. As long as inputs and outputs follow the same protocol, the upper-level VLM does not need to know which control method is used below.
Cross-room tasks create a problem. A two-dimensional image point can tell Pick what to grasp, but it cannot tell Navigate where a pill bottle seen minutes earlier is located. The G1 therefore explores the environment before executing tasks. During exploration, the system saves camera observations, D435i stereo data, LiDAR and SLAM information, robot joint poses and waypoints selected by Astra. Astra then participates in a Real2Sim process that organizes these observations into a digital environment in Isaac Sim. SLAM provides measured geometric constraints, visual data adds object and scene semantics, and joint states and waypoints rebind observations to the robot's own viewpoint.
Super Odometry obtains the G1's state in the real SLAM map, and Iterative Closest Point, or ICP, aligns the SLAM point cloud with the reconstructed environment in Isaac Sim. After registration, the real map and digital environment share spatial relations, and ego keyframes saved during exploration can be bound to that coordinate system. Spatial memory is split into semantic memory, which records what appeared in a keyframe, and metric geometry, which records where the robot was when the image was taken. For a medicine-fetching task, Astra can recover visual cues for a bottle or drawer from historical observations, then let Navigate return to the area recorded earlier. This is more stable than placing dozens of historical images directly into the VLM context.
Real2Sim also constrains scale. The report said that when reconstruction relies only on manually shot video, room dimensions must be estimated from visual appearance. Adding SLAM geometry collected by the robot constrains room layout and scale with measured point clouds, making the digital environment suitable for navigation and spatial reasoning. The spatial memory mainly solves the problem of re-finding a target after it leaves the field of view. When the robot reaches a drawer, the problem shrinks from meter-level navigation to centimeter-level manipulation, where a small three-dimensional error can cause the hand to miss the target.
For Pick, Astra provides a two-dimensional point in a 0 to 1000 range and specifies a hand. A segmentation module uses that point to cut the target from the background. Fast-FoundationStereo then computes depth from D435i stereo images. After obtaining depth for the target region, the system uses camera calibration to project pixels in the mask into three-dimensional geometry and generates a grasp pose analytically. For motion planning, the arm moves from its current pose to the grasp pose while the planner generates a spline reference and applies minimum-jerk timing constraints to keep velocity and acceleration changes smooth. The system solves inverse kinematics point by point along the trajectory and checks collision clearance for the entire swept motion, not only the endpoint.
Offline planning cannot handle visual drift during motion. As the humanoid approaches a table, its body moves and the head camera's view changes. A three-dimensional target that is off by only a few centimeters can cause a failed grasp. The system uses SAM 2.1 to track the target and SAMURAI's memory selection to maintain target state across frames, then applies visual servoing to continuously correct the approach direction from the new visual position. Small visual errors are corrected by the servo loop. If the gripper closes without establishing contact, Pick can switch to another grasp candidate or change the robot's stance and replan, within bounded attempts. Only when local recovery fails does the skill return a failure reason to Astra, which then decides whether to reposition, change target or adjust task order.
The model therefore does not enter every correction. It handles discrete task-state transitions, while local modules stabilize continuous states. Arm and hand control commands run at 250 Hz, AMO updates every five control ticks, or 50 Hz, and GPT Astra's remote inference operates at the second level. Keeping these scales separate means network jitter and model inference time do not directly affect body control, and pauses caused by the model occur mainly at skill switches.
Drawer operation extends the hierarchy to whole-body control. After the hand hooks the handle, moving only the arm backward would soon hit workspace and balance limits. The system makes the robot walk backward while maintaining the upper-limb target. AMO combines Sim2Real reinforcement learning and trajectory optimization, with one training goal being to let the G1 expand its operable space through the lower limbs and torso when upper-limb targets are outside the normal distribution. The technical chain then closes: the VLM outputs task-level goals, spatial memory turns historical vision into map positions, perception recovers three-dimensional geometry from two-dimensional targets, the planner generates executable trajectories, visual servoing absorbs local errors, and the whole-body controller keeps those actions stable on the real G1.
The system still has engineering limits. Real2Sim adds deployment preparation and API cost. Astra's remote inference creates pauses between skills. Local perception and planning currently require an RTX 4090 laptop, and heavier vision models would increase computing demands. Long tasks are limited by the robot arm's workspace, heat from hand servos and hardware durability. Skill coverage is a larger issue. Astra can rearrange Pick, Place, Navigate and Open Drawer, but it cannot generate contact dynamics that have never been implemented from language reasoning alone. If the skill library lacks table wiping, the robot cannot automatically obtain stable wiping control from a command. If there is no controller for deformable objects, upper-level planning cannot compensate for the missing lower-level capability.
The architecture does not require a single VLA to learn spatial memory, task decomposition, three-dimensional vision, grasping, collision avoidance and whole-body dynamics at once. Instead, it splits those problems into representations and control frequencies suited to each and connects them through explicit interfaces. The VLM sees skills and states, SLAM sees geometric space, vision models see pixels and depth, the motion planner sees trajectory constraints, and AMO sees the dynamics of the entire body. The humanoid thus begins to resemble a complex software system: high-level models make discrete decisions, a spatial system maintains external state, skills provide standardized call interfaces, and high-frequency control layers preserve physical stability. HomeBody's demonstrated capabilities remain limited, but its question is concrete. Competition in embodied intelligence may depend not only on who trains larger action models, but also on who designs better skill interfaces, more stable spatial state representations and clearer boundaries between high- and low-frequency control.