TL;DR
Figure AI's Helix 2.5 completed household chores in 30 unfamiliar Bay Area homes with zero prior training on those environments, hitting a 56% success rate versus 9% for a model without its human-behavior dataset.
What happened
- Figure AI unveiled Helix 2.5 on September 17, 2026, a zero-shot humanoid control model tested across 30 previously unseen Bay Area homes.
- 237 of 420 attempts succeeded across three full-body tasks: making beds (67%), folding towels (62%), and tidying toys (40%).
- Baseline comparison: the same hardware running a model trained from scratch, without Figure's dataset, managed only 9% on identical tasks.
- Hardware platform: Figure's 03 humanoid robot ran all trials; Helix 2.5 handles perception, decision-making, and physical control in a single network.
- Index dataset powers the leap: a pretraining corpus of human behavior recordings ingesting roughly 35 minutes of new human experience every second, publicly launched August 25, 2026.
Why it matters
- 6x improvement over baseline from pretraining alone, with no fine-tuning on target homes or objects, is a meaningful proof point for general-purpose robot learning.
- Zero-shot generalization at this scale (30 real, messy, varied homes) moves the goalposts: prior demonstrations typically relied on controlled or pre-mapped environments.
- Index as a strategic moat: the continuously growing human-behavior dataset is the engine behind the jump, making data flywheel velocity a core competitive variable in humanoid robotics.
- CEO Brett Adcock called Helix 2.5 the most important project Figure has undertaken, signaling the company views generalization, not hardware, as the defining race.
- 56% still fails nearly half the time, which is far below the reliability threshold for unsupervised home deployment, keeping commercial timelines uncertain.
What to watch next
- Whether Index's data ingestion rate translates to continued step-change improvements: the 35-minutes-per-second intake pace suggests Helix 3.x could arrive with substantially more training signal within months.
- Competitor responses from Physical Intelligence, 1X, and Agility Robotics: a 6x zero-shot generalization jump will pressure rivals to publish comparable real-world benchmarks or accelerate their own dataset strategies.
- The reliability threshold question: watch for Figure to announce a success-rate target (80% or higher is the informal bar cited in home-robotics research) as the trigger for commercial pilot programs.
Originally published on Present of AI, a daily source-linked AI news timeline. Read the full timeline or browse the open dataset.