Figure on Sept. 17, 2026, introduced Helix 2.5 and said its humanoid ran three long-horizon household behaviors across 30 Bay Area homes it had never seen, with no data collected in those houses. The company result matters because whole-body home work usually fails when the furniture, layout, and objects change; Figure is arguing that large-scale human-behavior pretraining can carry skills into unfamiliar rooms without on-site fine-tuning.
Zero-shot here does not mean the robot invented chores from nothing. Figure says the evaluation homes and the objects it touched were unseen, while the three behaviors were still specified with fine-tuning data collected elsewhere, and each task ran from a single fixed checkpoint with no weight updates in the test homes. The behaviors were living-room tidying, towel folding, and bed making.
Company claims: holding task-specific data, architecture, training, and evaluation fixed, Index pretraining raised full-task zero-shot success from 9% for a from-scratch policy to 56% for the Index-pretrained policy, more than six times higher. Figure says success required finishing the entire task, with no partial credit, and that a safety intervention counted as a failure. Separately, Figure says Index is generating about 35 minutes of new human experience per second and that it has committed $3.5 billion of compute to training Helix. It also reports an 8x nested Index-data scaling curve for held-out robot next-action prediction loss, with a largest-run forecast error of 0.54% of variation across that range; that curve is action-prediction loss, not household completion-rate scaling.
Independent coverage on Sept. 18 and 19 recounted the same disclosure with harder edges. Humanoid Analytics tallied 237 of 420 disclosed trials, including about 67% bed making (94/140), 62% towel folding (87/140), and 40% living-room tidying (56/140), and classified the evidence as Internal Testing rather than Operational Deployment. TechTimes framed the Index ablation as the technical heart of the release and noted that a 56% rate still means failure on roughly 44% of attempts under company criteria. Humanoid.guide warned the aggregate should not be read as uniform household competence. Gulf News called Helix 2.5 a research milestone rather than a consumer robot and said the 56% versus 9% comparison remains company figures not yet independently peer-reviewed.
What remains open is whether outside labs will reproduce the 30-home grades, how often safety stops appear in longer unsupervised runs, and whether next-action loss scaling turns into higher completion rates on messier homes. Company eval design and third-party caveats should be weighed together until independent replication lands.