Homebody: Stanford pilots a humanoid with GPT-6 Astra and zero learned policy — 7/100 on the desk tasks benchmark
🔎 A humanoid without a learned policy, and why September 2026 matters
A Unitree G1 just tidied up a kitchen it was discovering for the first time. No prior teleoperation, no policy trained for weeks on thousands of demonstrations: GPT-6 Astra, OpenAI's frontier VLM, directly calls the robot's skills the way a developer calls functions. The project is called HomeBody, it comes out of Stanford's Movement Lab and Caltech, and it was published in late September 2026.
What stands out is the VLA (Vision-Language-Action) layer: the "foundation" robotics standard that required collecting demonstrations, training a specialized controller, and then hoping it would generalize. HomeBody replaces this trained layer with a direct dialogue between the model and a skill library. The demo is public on YouTube — and hard to dismiss as vaporware: you can see the robot navigate, open drawers, grasp and place objects.
The timing says everything. On September 12, 2026, UBTECH inaugurated its Liuzhou factory, capable of assembling one humanoid every ten minutes. Three weeks later, Stanford demonstrated that an LLM could drive these machines without any environment-specific training. The body is industrializing while software changes paradigm. One number remains sobering: 7/100 on the desk task benchmark. Both facts are true at the same time — and it is precisely this tension that needs decoding.
Key Takeaways
- Zero learned policy: HomeBody (Stanford TML + Caltech, late September 2026) removes the VLA layer. GPT-6 Astra pilots a Unitree G1 by directly calling 5 composable skills: navigation, pick, place, drawer opening, pick in drawer.
- Spatial memory before action: the robot first explores the room, builds a Real2Sim digital twin in Nvidia Isaac Sim from camera streams, SLAM and joint data, then logs objects and their locations.
- A serious perception stack: SAM 2.1 for segmentation, Samurai for tracking, fast foundation stereo for depth — and arm commands at 250 Hz.
- Two-sided results: a never-seen kitchen cleaned end to end, but 7/100 on Stationary Bench (100 office tasks, 200 trials) vs 0 for the specialist Malmo Act 2 — with a median progress score of 46/100.
- Industrial context: the Unitree G1 body is now mass-produced (~$13,500, September 2026, check unitree.com), and UBTECH's Liuzhou factory is targeting more than 10,000 humanoids per year.
Recommended Tools
| Tool | Main use | Price (September 2026) | Best for |
|---|---|---|---|
| HomeBody (Stanford TML) | Project page: architecture, demos, results | Free (research) | Following and understanding the project |
| GPT-6 Astra (OpenAI) | Piloting VLM, structured tool calls | API, pricing n/a | Orchestrating robotic skills |
| Unitree G1 | Humanoid body used in the demo | ~$13,500 (check unitree.com) | Reproducing a humanoid setup |
| Nvidia Isaac Sim | Real2Sim digital twin | Free (download) | Simulation and spatial memory |
| SAM 2.1 (Meta) | Object segmentation in video streams | Open source, free | Scene perception |
A VLM instead of a policy: the HomeBody architecture in plain terms
HomeBody removes the trained controller that traditionally separated the language model from the robot. In its place, GPT-6 Astra runs remotely and selects skills via structured tool calls, drawing on the ego view, the map, the gripper state, and observations recalled from spatial memory.
A recap of the old-school standard. A VLA system is built in three steps: collect demonstrations — often teleoperated by hand — train a policy that maps perception to action, then deploy with fingers crossed. Every new task domain restarts the data collection machine.
It's slow, expensive, and generalization remains the chronic weak point of the paradigm. As The Decoder sums it up (September 27, 2026), HomeBody outright removes that trained control layer between the model and the robot.
The approach inverts the logic. The 5 skills — navigation, pick, place, drawer opening, pick from drawer — are reliable primitives, written and tested once and for all. Generalization no longer comes from a policy: it comes from the VLM, which reasons about a novel scene and composes the primitives. The model is swappable: replace GPT-6 Astra with another frontier VLM, and the architecture still holds.
The pipeline: explore, memorize, act
First phase, exploration. Dropped into an unknown kitchen, the robot roams the room and builds a Real2Sim digital twin in Nvidia Isaac Sim, from camera, SLAM, and joint data. Objects and locations are logged into a spatial memory — the robot is literally building its own mental map.
Second phase, execution. The prompt is deliberately under-specified — "tidy up the kitchen" — and it's the model that breaks it down into long sequences of tool calls. No environment-specific training: that's the point AI Weekly highlights in its September 27, 2026 analysis.
Concretely, the decision loop resembles that of a coding agent. The model receives a structured context — ego view, map, gripper state, recalled observations — and returns a call: navigate to X, pick Y, open drawer Z. The skill runs locally, then reports back. The VLM doesn't drive the motors: it drives intentions.
On the sensor side, the stack is classic but serious: SAM 2.1 segments the objects, Samurai handles their tracking in the video stream, fast foundation stereo provides depth. Arm commands go down at 250 Hz — the frequency that makes grasps stable on a humanoid.
My take, after watching the demo several times: the genius of the project isn't in the prompt, it's in the interface. A clean, documented skill library, with states returned to the model — that's exactly the discipline you'd demand of a good API. Robotics has just become a software integration problem.
Kitchen cleaned, office at 7/100: the results, unvarnished
On the flagship demo, HomeBody cleaned a never-before-seen kitchen, drawers included, from underspecified prompts. On the office benchmark, it caps out at 7/100. The two numbers tell the same story: the paradigm works, reliability hasn't caught up yet.
What "cleaning a kitchen" actually means
Navigating a never-before-seen room, identifying misplaced objects, opening drawers, extracting their contents, putting each item back — all chained together across a long horizon of tasks, with no prior training on this kitchen. Developments Today confirms the setup: full autonomy in the unknown environment, retrieving objects from drawers included.
Stationary Bench is the test that stings: 100 office tasks, 200 trials. GPT-6 Astra via HomeBody scores 7/100. The specialist Malmo Act 2, a policy trained for this type of task, scores 0. And the median progress stands at 46/100: on half the tasks, the robot makes it halfway before failing.
| Model | Approach | Score (100 tasks, 200 trials) | Median progress |
|---|---|---|---|
| GPT-6 Astra (via HomeBody) | Frontier VLM + skill library | 7/100 | 46/100 |
| Malmo Act 2 | Trained specialist policy | 0/100 | — |
Reading these numbers calls for rigor. A 7/100 against a 0/100 doesn't mean "GPT-6 Astra is bad": it means "the specialist produces nothing relevant outside its training domain, the generalist at least does something". It's the same pattern as in code, where FrontierCode, the benchmark from Cognition that buries SWE-Bench, showed that frontier models crush specialists as soon as you measure the actual quality of pull requests.
But a 93% failure rate on trivial tasks — placing an object, opening a drawer — remains an acknowledged plateau. The median progress of 46/100 suggests failures late in the chain: fine-grained perception, contact precision, handling the unexpected. The paradigm has won the architecture match; it hasn't yet won the reliability match.
Why removing the VLA layer is the real story
Because that layer stacked two locks on top of each other: a data bottleneck and a generalization ceiling. HomeBody doesn't improve it — it bypasses it, and it's that bypass that is strategic.
| Classic VLA pipeline | HomeBody | |
|---|---|---|
| Data | Teleoperated demonstrations to collect | No specific collection |
| Training | Policy per task domain | None |
| Generalization | Limited to the seen domain | Never-seen kitchen, out of the box |
| Update | Full retraining | VLM swap |
| Marginal cost of a new task | Weeks | One prompt + existing skills |
The key point is the VLM's swappable nature. Robotics thus directly inherits the LLM progression curve — the very one that has seen GPT-5.5 reign at 98.2 on our agentic index since June 2025, and that allowed Claude to break the theoretical physics computing record: the Yang-Mills amplitude at 9 loops for $2,000 of compute. Every frontier model jump carries over to the robot without touching a single line of the control stack.
Important nuance: the VLA isn't dead. It goes back to being one component among others — possibly just another skill in the library. What HomeBody kills is the monopoly of the trained policy as the only path to action. Teams that have invested months in demonstration collection would do well to reevaluate their roadmap.
The body now costs almost nothing: software becomes the front line
Hardware is no longer the bottleneck. The Unitree G1 from the demo is in serial production at around $13,500 (September 2026, check unitree.com) — the price of a good used car for a dual-arm humanoid.
On September 12, 2026, UBTECH opened its Liuzhou factory: a rate of one humanoid every ten minutes, i.e., more than 10,000 units per year. As we detailed in our article on Automate 2026 and the reality check of humanoid robotics, China has already moved on to mass production while the West remains at the pilot stage.
The conjunction of the two September events is the real signal. When the body drops to $13,500 and the brain becomes an API call, the barrier to entry for robotics research collapses: a university lab, or even a three-person team, can reproduce a HomeBody. The monopoly of the big robotics labs on this kind of result is over.
What This Changes for Developers
To a developer, HomeBody looks strikingly familiar: an agent that calls tools. Robotics converges with code — you write reliable primitives, document their inputs and outputs, and let a frontier model orchestrate.
This is exactly the thesis of Feather, the startup building the Android of robotics for developers: a common abstraction layer on top of heterogeneous hardware, where skills are written as reusable modules. HomeBody is the academic proof; Feather is the commercial attempt.
Second parallel: benchmarks. In code, FrontierCode replaced SWE-Bench because it measures the real quality of merged pull requests, not synthetic patches. In robotics, Stationary Bench plays the same role: 100 real tasks, 200 trials, zero leniency. Both movements say the same thing — the era of synthetic benchmarks is over.
The sector's economics are shifting along with it. When value migrates from the policy to the skills, a well-written "drawer opening" primitive becomes a reusable asset across dozens of tasks and multiple generations of models. Good news for developers; less good for those who bet their valuation on proprietary demonstration data.
If you want to get started, the order of priorities is clear: robust skills first, then a solid orchestration model. Our guide to the best LLMs for coding serves as a starting point for choosing the latter — agentic capabilities now matter more than raw chat scores.
The limits: 93% failure, cloud latency, and benchmarks to take with a grain of salt
The number that shouldn't get buried under all the enthusiasm: 7 successes out of 100 desk tasks. The performance plateau of frontier LLMs in robotics still stands at 93% failure on tasks a five-year-old breezes through without a second thought.
First technical limitation: GPT-6 Astra runs remotely. Every decision goes through the cloud, which means latency, network dependency, and privacy concerns — your kitchens become 3D scenes sent off to your model provider. Skills, on the other hand, execute locally: hence the importance of robust primitives, capable of securing the robot's state if the connection drops.
Second methodological limitation: caution with the scores. GPT-5.6 Sol on Cerebras at 750 tokens/second — and the benchmark gaming trap that METR just uncovered is a reminder that an isolated number can be optimized for its own sake. Stationary Bench has the merit of being filmed — the trials are public on YouTube — but robotics doesn't yet have its own METR to audit benchmarks.
Third limitation: safety. A model that selects physical actions based on recalled observations must prove it isn't recalling just anything. Stale spatial memory — an object moved between two sessions — becomes a physical risk, not just a display bug. It's an entire research program in its own right, and there's no standard answer to it today.
❌ Common Mistakes
Mistake 1: Believing the LLM "sees" the world
The VLM does not replace perception: SAM 2.1, Samurai, and stereo vision do the work of segmentation, tracking, and depth. Without this stack, the model receives raw pixels and fails. Solution: treat perception as a first-class investment, not as an implementation detail.
Mistake 2: Underestimating the skill library
HomeBody's 5 skills are engineered, tested primitives, with states that the model can exploit. Remove that rigor and "plug a GPT into an arm" produces nothing usable. Solution: treat each skill as a product — tests, documentation, explicit failure modes.
Mistake 3: Reading 7/100 as a model failure
The Malmo Act 2 specialist scores 0, and the median progression is 46/100. The absolute score says less than the relative gap and the partial progression. Solution: read a robotics benchmark through comparison and video — the trials are filmed — never as an isolated number.
Mistake 4: Ignoring that the brain is in the cloud
Remote inference adds latency and network dependency to a physical control loop. Solution: plan for local skills capable of putting the robot into a safe state if the connection drops, and test this degradation explicitly — not just the nominal case.
❓ Frequently Asked Questions
What exactly is HomeBody?
A research project from Stanford and Caltech's Movement Lab, published in late September 2026. It plugs a frontier VLM — GPT-6 Astra — directly into a library of 5 robotic skills, without the usual trained VLA control layer. The robot first explores the room, builds a digital twin, then acts.
Why remove the VLA layer?
Because it requires collecting demonstrations and training a policy for every task domain. HomeBody shifts generalization to the VLM, which composes reliable primitives via structured tool calls. The result: zero environment-specific training, and a swappable model that benefits directly from the progress of frontier LLMs.
How much does the demo robot cost?
The Unitree G1 sells for around $13,500 (September 2026, check unitree.com). It's a mass-produced humanoid with grasping arms. Add to that the VLM's inference costs, since it runs remotely, and above all the engineering time for the skill library — the main cost item in practice.
Can GPT-6 Astra be replaced with another model?
Yes, it's designed that way: the VLM is swappable. The model receives the ego view, the map, the gripper state, and recalled observations, and returns structured tool calls. Any frontier VLM that's solid at agentic work can claim the role — including Gemini 3.1 Pro or Claude Opus 4.7.
Are teleoperated demonstrations still needed?
For the skills, human engineering is still required: HomeBody's primitives are written and tested, not learned through imitation. What disappears is the massive per-task collection of demonstrations. The cost/value ratio flips: a few weeks of engineering per skill, versus months of teleoperation per domain.
Can HomeBody tidy up your kitchen tomorrow?
No. 7/100 on trivial tabletop tasks, mandatory cloud inference, and no safety certification: it's a research prototype. What changes is the direction — training-free generalization now exists. Industrial reliability still lies ahead, probably a few model generations away.
✅ Conclusion
HomeBody skips the VLA layer and proves that a frontier VLM can pilot an off-the-shelf humanoid without any task-specific training — but its 7/100 on Stationary Bench is a reminder that the reliability plateau still sits at 93% failure on trivial tasks. The paradigm flipped this month; reliability will flip one model iteration at a time. To dig deeper, the HomeBody project page brings together architecture, demos, and results — and the YouTube video says more than all the diagrams.