Figure AI's Helix 2.5 Reaches 56% Zero-Shot Success in 30 Unseen Homes
Figure AI said Helix 2.5 achieved a 56% zero-shot success rate in 30 unseen Bay Area homes, versus 9% without Index pretraining.
In the trials, Helix 2.5 folded towels, made beds and picked up objects. It succeeded 237 times out of 420, for 56%. The evaluation was blind, and no single tested task accounted for more than 1.90% of Index pretraining data, Figure said. The company isolated the toys, towels and bedding used in the tests, screening them first with an AI model and then by humans to confirm they had not appeared in task training data; the 30 homes kept their existing sofas, beds and folding surfaces.
Figure compared the result with the same hardware and task data but without Index pretraining. That policy succeeded 9% of the time. Index pretraining raised zero-shot success from 9% to 56% and lifted overall success by 522%, according to the company. Figure also compared Helix 2.5 with Helix 02, its previous model, which needs data collected at the target site; Helix 2.5 used half the adaptation data and matched Helix 02's single-site success rate across the 30 unseen homes.
CEO Brett Adcock called it 'the most important project we have ever done,' and AI director Corey Lynch said the robot recorded success in every one of the 30 homes in the release video. The report said the result contrasts with many prior home-robot demonstrations, including Tesla's Optimus and 1X's Neo, which rely heavily on site-specific data or human remote operation.
The change rests on Index, which Figure describes as its largest and most diverse robot training dataset. Figure quietly launched a data-collection app in May and formally introduced Index on Aug. 25. It said Index had 264,000 downloads, covered 108 countries, had more than 44,000 weekly active creators, and had received more than 16 million uploaded videos. Upload speed reached 30 minutes of new video per second, equivalent to 4.9 years of human work entering the system each day, and later rose to 35 minutes per second, according to the report.
Figure has paid creators $15 million and plans to spend more than $1 billion over the next 12 months on data and compute, aiming to expand collection 100-fold. Each 1,000 hours of collected data contain 373 unique tasks, 1,146 unique manipulated objects and 116 unique environments, the company said. The data pass through five steps, filtering, anti-cheat checks, deduplication, distribution rebalancing and text labeling, before training.
Figure said doubling Index pretraining data repeatedly reduced action-prediction error in a predictable way when model scale and downstream training conditions were fixed. The experiment used four data scales, with the largest eight times the smallest; the error curve was smooth enough to predict the largest training loss from small-scale experiments, with a deviation of only 0.54% of the overall fluctuation range. The company calls this a scaling law for human-to-humanoid transfer.
Other companies are also trying to widen their sources of experience. 1X's Neo opened for home preorders in February at $20,000, but most complex actions, such as taking water from a refrigerator, loading dishes into a dishwasher or folding a sweater, require real-time takeover by company remote operators code-named 'Turing' using VR equipment. 1X has fed those teleoperation data and home videos into a world model released in January, hoping to reduce dependence on human takeover. Tesla's Optimus relies more on large-scale human teleoperation, with operators using sensor-equipped devices for tasks such as sorting and using power tools, then giving the trajectories to a neural network to imitate and optimize.
Figure's approach is to have the model learn general behavior patterns from massive human videos, then adapt to specific actions with a small amount of task data. In theory, that avoids sending workers to collect data or take remote control in each new home.
The test still has limits. It covered only three tasks, and the report's author said more tasks, such as wiping mirrors, organizing food and cooking, would be preferable. If Figure can expand to 10 to 12 tasks and raise zero-shot success to 80% to 90%, deployment in homes without first collecting that home's data would be closer. Figure says the 56% success rate does not mean the robot will always fail, because Helix 2.5 can self-correct and retry in long-horizon tasks. Rented homes provide different layouts but may not cover the full disorder of daily life, such as Lego bricks on carpet, charging cables in sofa seams or toys a child has just piled up, all of which could change a route the robot could otherwise pass. Cleaning up before the robot arrives would defeat the purpose for many buyers.
Figure is preparing more resources for training. In early September, cloud company Nscale announced it would provide at least $3.5 billion in compute to Figure, with the two sides planning to expand that to more than $6 billion and deploy up to 100,000 Nvidia Vera Rubin chips; the first equipment is expected to land in Bastrop, Texas, in the second half of 2027. As part of the deal, Nscale took an equity stake in Figure and became its preferred compute supplier. Adcock said the biggest bottleneck to putting humanoid robots in every home is data and compute.
Index is already collecting human experience on a large scale. Figure's report indicates that transfer improves predictably as data grows. The question is how high 56% can be pushed after several larger Index training runs and a larger task set.