📑 Table of contents

Spirit AI: the "GPT-3 moment" for robot brains is expected in mid-2027 — provided dirty data is accepted

Deep Tech 🟢 Beginner ⏱️ 15 min read 📅 2026-09-18

Spirit AI: the "GPT-3 moment" for robot brains is expected by mid-2027 — provided dirty data is accepted

Chinese humanoids have won the image race: they sprint, dance, and chain backflips on command. On September 18, 2026, Reuters published an interview with Gao Yang, co-founder and chief scientist of Spirit AI, and moved the front line one notch further: "The brain is indeed the weakest link in the complete robotics stack."

His prediction comes with a date: a "GPT-3.0" equivalent — the model that made ChatGPT possible — for robot brains by mid-2027. His counter-promise is just as specific: no humanoids in homes for at least another eight years, i.e., ~2034.

The most interesting part is neither the optimistic date nor the pessimistic one. It's the method: roughly 1,000 contractors equipped with sensors repeating mundane gestures in collection centers in Beijing, an openly embraced preference for "dirty data" over clean data, and a head-on bet against simulators.


Key Takeaways

  • Gao Yang (Spirit AI, robotics professor at Tsinghua University) announces a breakthrough in humanoid "brains" by mid-2027, comparable to the impact of GPT-3.0 that made ChatGPT possible (Reuters, September 18, 2026).
  • The bottleneck is no longer hardware but real-world motor data: Spirit AI is mobilizing ~1,000 contractors equipped with portable capture gear, at home or on the factory floor.
  • "Dirty data" — a diverse range of imperfect gestures — advances its models faster than the 50 repetitions required elsewhere in China for a single "clean" movement. The opposite of LLM pipelines.
  • Field evidence: a 90% success rate on simple tasks in a structured living-room setting, and dozens of Moz1 (wheeled humanoids) already deployed at CATL and JD.com.
  • Full timeline: an industrial window over the next 1 to 2 years, simple commercial services by ~2028, homes not before ~2034.
  • A deliberate technical bet: simulators fail on deformable bodies (electrical cables), hence the massive data collection in real-world conditions.

The Players Behind the "Robot Brain": Three Clocks, One Window

Three independent Chinese players are converging on 2027 — and this convergence is worth more than any isolated prediction.

Player Approach Stated deadline Field signal
Spirit AI (Gao Yang) Real-world "dirty" data, ~1,000 contractors Mid-2027 Dozens of Moz1 units at CATL and JD.com
ACE Robotics Embodied AI End of 2027 Public announcement by the chairman
Unitree (Wang Xingxing) Humanoids, hardware-first "Two to three years at the earliest" Stock market listing, +460% on day one

When competitors who are not coordinating with one another agree on the same window, it's rarely marketing: it means the roadmap is dictated by a measurable variable — here, the accumulation of motor data. The financial context also weighs heavily on this reading: Unitree went public with +460% on day one and a ~$50B valuation — a signal that the market is already paying for the "ChatGPT moment" narrative for robots (we detailed its implications here).

To situate this micro-sector within the broader AI landscape, our overview of current AI trends is updated every month.


Who is Gao Yang — and why his dates carry weight

He's not a tech influencer; he's an academic with robots in real-world production — statistically, the worst possible profile for overselling a timeline.

Gao Yang is an assistant professor of robotics at Tsinghua University, one of China's most powerful academic hubs on the subject, and co-founder/chief scientist of Spirit AI, a company of around 300 people. The startup has raised more than 2.9 billion — the currency isn't specified in the Reuters excerpt, most likely yuan (check the original source) — and Gao declines to comment on a potential IPO.

His definition of the milestone is remarkably precise: "We anticipate reaching the GPT-3.0 milestone by mid-2027. You will be able to speak to a robot in natural language, and it will execute a series of reasonable physical actions to attempt the task."

To my mind, this is what sets this prediction apart from the usual Silicon Valley noise: a date, a testable criterion, and a symmetric admission (houses are eight years away). A founder selling a five-year hope AND a five-year disappointment is a founder who has done the math.


The bottleneck has moved: from metal to motion data

Hardware has surpassed what software knows how to exploit — robots' limiting factor is no longer the mechanics, it's the training.

That's the very framing of the Reuters/CP24 recap: after impressive hardware progress, Chinese robotics firms are shifting the focus toward the software that determines robots' intelligence and economic productivity — "embodied AI."

The scene described by a Reuters witness at Spirit AI's Beijing offices sums it all up: dozens of young people equipped with sensors repeat gestures on loop — opening fridges, unlocking safes, chopping vegetables with a knife. Nationwide, roughly 1,000 contractors, equipped with portable capture gear, supply human motion data from their homes or from factories.

A detail worth pausing on: some contractors collect data from their own homes. That's not about saving on expenses — it's data distribution engineering. If robots will one day have to operate in living rooms, you might as well train the models in real living rooms, with their rugs, their sofas, and their imperfect obstacles.

The economic read: for LLMs, text data was free and had already been collected by all of humanity. Here, every hour of movement must be paid for, body by body, hour by hour. Motion data has become a human-capital asset — Spirit AI's ~1,000 contractors are, in robotics, the equivalent of the text labs' GPU farms.

And this "physical" pivot of AI didn't come out of nowhere: we documented it in the September 17 Veille, with the shipping figures confirming the sector's physical turn. Spirit AI is its most explicit operational translation to date.


"Dirty data": imperfection beats clean data

The most counterintuitive lesson from the interview comes down to one sentence: the variety of imperfect movements advances models faster than repeated perfection.

The point of comparison comes from Chinese ground itself: in other training facilities, it takes more than 50 repetitions to obtain a single "clean" movement with the required precision. Spirit AI found that "dirty data" — a more diverse range of gestures — moves its models forward faster.

My analytical hypothesis: a movement captured 50 times by the same body, from the same angle, in the same room, doesn't learn a task — it learns one person's gesture. Dirty variety covers the real distribution of executions: different body types, varied angles, hesitations, trajectory corrections. This is, roughly speaking, regularization through diversity applied to motor skills.

The apparent paradox with LLMs — where data gets cleaned, deduplicated and filtered — is less crazy than it seems. In text too, it's the diversity of sources, not their polishing, that has historically driven generalization. Cleaning serves readability, not learning.

The field is in fact exploring every data strategy in parallel, and they are almost opposites. OM-1 from Reward AI, a robotic policy trained solely on demonstrations bets on pure demonstration; Skild S1 learns a ten-minute task from a single video pushes extreme efficiency. Spirit AI takes the opposite path: volume, variety, impurity. My bet: the two schools will converge — efficiency for rapid adaptation, dirty volume for coverage — but to hold to a 2027 timeline, today it's the dirty approach that wins.


Why Spirit AI refuses the simulator shortcut

Because simulators are very good at simulating… rigid objects. But the real world is full of soft objects.

"Simulators handle rigid bodies well, but flexible objects like deformable electric cables remain a problem," sums up Gao Yang — precisely where many competitors are betting heavily on simulation to reduce training costs.

Take the example they chose, because it's perfect: an electric cable. It gets plugged in, unplugged, pulled, tangled, jammed — in factories just as much as in homes. If simulation fails on deformable bodies, a policy trained in a simulator will fail exactly at the most mundane failure points of the physical world. You'll end up with a robot that's brilliant with geometric objects and falls flat on its face at the first carelessly stashed charger.

The cost of the real-world bet is considerable: ~1,000 contractors, portable equipment, dedicated centers in Beijing. This isn't a method, it's an infrastructure — and the raise of over 2.9 billion should be read as the price of that infrastructure, not as a classic R&D budget.

My take: as long as simulation engines can't handle deformables, real-world data remains a defensible moat, perhaps the only durable one in the sector. But it's a time-limited moat — the day simulation catches up on this specific point, the advantage of the big data collectors will flip fast. That's the variable I'd be watching before drawing any definitive conclusion about who "wins" 2027.


Dozens of Moz1 already in production at CATL and JD.com

The 2027 promise doesn't start from a mockup: Spirit AI's robots are already working on real production lines, for demanding customers.

The facts: dozens of Moz1 — Spirit AI's wheeled humanoids — are deployed on the lines of battery manufacturer CATL and retailer JD.com. A detail that counts double: JD.com is also an investor in the startup.

Two readings of this "wheeled" choice. First, engineering: on a factory floor, the stability and cost of a wheeled base still beat bipedal walking, whose only current added value is... doing backflips in demos. Then, the capital structure: a customer-investor locks in both captive demand and a product feedback loop — the Moz1 learn on tasks paid for by real customers, not on test benches.

This is exactly what Gao Yang's roadmap promises: the next one to two years constitute "the initial window for industrial applications". Structured industrial environments are where the 90% success rate measured by Spirit AI translates best — living rooms will come later, and for good reason.

Last point, the most strategic in my view: dozens of robots coordinated on a single line is no longer one robot, it's a physical multi-agent system. We argued in our article on Agentic AI for robotics that the next "ChatGPT moment" for robots would be collective, not individual. The CATL/JD.com deployment is its first foreshadowing at scale.


2027 promised, ~2034 for the living room: what the timeline actually says

Gao Yang is selling a capability breakthrough, not domesticity — and the gap between the two is precisely what most commentary glosses over.

Timeline Milestone Source
Mid-2027 "GPT-3.0 milestone": natural language → reasonable physical actions MarketScreener, 18/09/2026
2026-2028 Initial window for industrial applications Reuters, 18/09/2026
~2028 Robots in commercial service settings, simple tasks Reuters, 18/09/2026
≥ 8 years Entry into homes (~2034) Reuters, 18/09/2026

The full roadmap quote is worth reading: "The next one to two years mark the initial window for industrial applications. Two years from now, we'll see robots deployed in commercial service settings doing simpler tasks. Entering homes is far harder than both."

Why homes are a different sport altogether: fine motor skills remain an open challenge — unscrewing a bottle cap is called out by name — and handling never-before-seen tasks is unsolved. Spirit AI's 90% success rate is measured on simple tasks, in structured living rooms. A "lab" living room, in other words. The gap between that number and your actual apartment is the entire road that remains.

And even the industry's most aggressive predictions aren't about domesticity: ACE Robotics' chairman points to a "ChatGPT moment" by the end of 2027, while Wang Xingxing, founder of Unitree, sticks to "two to three years at the earliest" (Reuters, August 21, 2026). No credible voice is promising a butler by 2027. The realm of the possible for 2027-2028 is industrial and commercial, nothing else.


The GPT-3 parallel: what it foretells, what it doesn't promise

The parallel is accurate on capability dynamics, misleading on economics: motor data has no free web equivalent.

What the parallel gets right: GPT-3 wasn't a product, it was a capability milestone that made a product possible. If mid-2027 delivers robots that "listen in natural language and attempt a sequence of reasonable actions," we'll have the same structure — a raw capability, still unusable at scale, but one that changes the slope of the curve. That's exactly what GPT-3 was for ChatGPT.

What it doesn't say is the economic asymmetry: LLMs grew on web text, collected once, endlessly reusable, at near-zero marginal cost. Here, every hour of human gesture costs hours of paid labor. Scaling laws may exist for motor skills — Spirit AI believes in them enough to pay ~1,000 contractors — but the economic unit has no precedent.

There is also a temporal asymmetry that the narrative glosses over: when robotic brains reach their "GPT-3," text models will be far ahead. The 2026 leaders in agentic benchmarks are named GPT-5.5, Gemini 3 Pro Deep Think, or Claude Opus 4.7 — see our Claude, GPT, Gemini, Llama comparison for 2026 and our selection of the best LLMs for coding. Robotics will climb, years behind, a curve that text has already descended, and "embedded" intelligence will inherit foundation models that are already highly advanced — which is, incidentally, its opportunity.

The real lesson of 2022, finally: the "ChatGPT moment" didn't come from a benchmark, but from an interface that made an existing capability suddenly useful. If natural-language interaction becomes reliable on a robot, the moment may arrive faster than skeptics think. That is precisely the bet — no more, no less.


❌ Common Mistakes

Mistake 1: reading "GPT-3 moment" as "household robots in 2027"

The announced breakthrough concerns motor generalization, not your kitchen. Homes are at least eight years away according to Gao Yang himself — the man giving the optimistic date. The real 2026-2028 opportunity is industrial and commercial; whoever waits for the butler misses the only investable window.

Mistake 2: overselling the 90% success rate

This figure covers simple tasks, in structured living rooms — not a real apartment, let alone a task never seen before. Fine motor skills (unscrewing a cap) and out-of-distribution generalization remain the walls. Never build a business case on "90%" without those two qualifiers.

Mistake 3: applying LLM economics to robotics

"Data will eventually become abundant and free": that was true for text, it is false for movement. The strategic asset is not the algorithm but the collection infrastructure — here, ~1,000 contractors and wearable equipment. Reasoning "compute + scraping" in robotics means getting the marginal cost wrong, and therefore the business model.


❓ Frequently Asked Questions

What is the "GPT-3 moment" for robots?

It's a capability milestone: a robot you can speak to in natural language and that executes "a series of reasonable physical actions" to attempt the task, even imperfectly. Just as GPT-3 made ChatGPT possible without being a finished product, this milestone would make useful robots possible. Gao Yang places it in mid-2027.

When will humanoids enter our homes?

Not before ~2034: Gao Yang speaks of at least eight years after the breakthrough. The home stacks up the worst difficulties — unstructured environment, unforeseen objects, fine motor control. Even ACE Robotics, the industry's most aggressive prediction, evokes a ChatGPT moment in late 2027, not domestic adoption.

Why would "dirty data" work better than clean data?

Because it covers the true distribution of movements: body types, angles, rhythms and hesitations all vary. Where other centers require more than 50 repetitions for a "clean" movement, Spirit AI finds that this imperfect diversity generalizes better. A gesture repeated fifty times in exactly the same way teaches the robot one person, not a task.

Why does Spirit AI avoid simulators?

Simulators handle rigid bodies well but fail on deformable objects — electrical cables first and foremost, omnipresent in factories as much as in homes. A policy trained in simulation risks failing precisely on the most mundane cases. Spirit AI prefers to pay for real-world data collection (~1,000 contractors) rather than take that risk.

Can you invest in Spirit AI?

Not on the stock market for now: the company is private, and Gao Yang refuses any comment on a potential IPO. Listed exposure to the sector notably runs through Unitree, which went public with a +460% first-day gain at a ~$50B valuation — with the volatility that this kind of listing implies.


✅ Conclusion

The humanoid race has shifted from hardware factories to data collection rooms — and it's Spirit AI's recipe (real, dirty, expensive) that now sets the timeline: capabilities by mid-2027, factories first, homes around 2034. To follow this physical shift of AI month by month, head over to our current AI trends.