📑 Table of contents

**Figure Index: Figure AI transforms human videos into training data — 264,000 downloads, 16 million videos, and $15M paid to contributors**

Skynet Watch 🟢 Beginner ⏱️ 13 min read 📅 2026-10-05

Figure Index: Figure AI turns human videos into training data — 264,000 downloads, 16 million videos, and $15M paid to contributors

🔎 The humanoid data flywheel hits industrial scale

Figure AI isn't just building robots. The company is buying human motion at scale through Index, a mobile app that pays contributors to film their everyday tasks. In just a few months: 264,000 downloads across 108 countries, 16 million videos uploaded, 44,000 weekly active users, and $15 million paid out. Behind these company-reported figures lies an infrastructure that ingests the equivalent of 4.9 years of human work per day. The goal: train Helix, the visuomotor control model designed to enable Figure 02 humanoids to manipulate any object in any environment. The bet is clear: whoever owns the largest dataset of real-world manipulation wins the autonomy race.


The essentials

  • Index isn't your classic crowdsourcing app: it's a proprietary data factory for training Helix, Figure AI's visuomotor model.
  • Key figures (company-reported, May 2026): 264k downloads, 108 countries, 16M+ videos, 44k+ WAU (then 69k+), $15M paid to creators, $1B+ committed over 12 months for data + compute.
  • Critical infrastructure: the app processes ~30 minutes of video per second (≈4.9 human years/day), requiring a complete overhaul of the data stack (24/7 availability, continuous compute, real-time feedback, embedding-based deduplication).
  • Business model: Figure buys raw data instead of scraping it — a major legal and ethical difference from LLM training practices.

Tool Primary use Price (May 2026) Best for
Figure Index (iOS/Android) Paid collection of human manipulation videos Free (compensation per validated video) Contributors looking to monetize their everyday gestures
Hostinger Data infrastructure hosting / robotics showcase sites From €2.99/month (May 2026, check hostinger.com) Robotics startups building their video ingestion pipeline
Roboflow Annotation, vision dataset management, ML pipelines Free up to 10k images; paid beyond that (May 2026) ML teams curating Index videos for Helix training
Weights & Biases Experiment tracking, model versioning, visualization Free for personal use; teams from $200/month (May 2026) Research on Helix-type visuomotor models

Index's architecture: a data factory, not just another app

Index is not your typical consumer app. Coming out of stealth (formerly Project Go-Big) with a public iOS/Android rollout, it was designed from day one as a high-throughput data acquisition system for training Helix. The pipeline is brutal: a contributor wears a sensor-equipped headset, films a household task — making a bed, folding laundry, putting away dishes — and the upload is ingested, deduplicated, audited by human analysts at the user level, then stored for continuous training.

The infrastructure has to absorb 30 minutes of video per second. That's roughly 4.9 years of human labor uploaded every day. Figure AI had to rebuild its entire data stack: 24/7 availability, continuous compute, real-time feedback to contributors, embedding-based deduplication with a similarity threshold to avoid redundancy. This isn't classic mobile app engineering — it's web-scale data platform engineering.

The distinction is crucial: the 16 million videos announced are raw uploads. The fraction actually retained for training Helix — after deduplication, quality control, and human auditing — is significantly lower. Figure does not disclose this retention rate. In tech journalism, you separate vanity metrics (total uploads) from utility metrics (hours of unique, diverse, annotated data that actually improve the model).


Helix: the visuomotor model that justifies the investment

Helix is the "brain" that Figure AI trains for its Figure 02 humanoids. It's a visuomotor control model — it takes video (onboard cameras) and proprioception as input, and outputs continuous motor commands for manipulation. No modular perception → planning → control pipeline: Helix learns the whole thing end-to-end, from pixel to motor torque.

This is where the Index dataset becomes critical. Visuomotor models are hungry for diversity: thousands of objects, textures, lighting conditions, spatial configurations, grasping strategies. Web scraping doesn't provide that. Academic datasets (EPIC-KITCHENS, Something-Something) are too small, too scripted, too clean. Figure needs dirty, real, diverse data — exactly what ordinary humans produce in their kitchens, living rooms, garages.

The $1 billion commitment over 12 months for data acquisition and Helix compute signals the absolute priority. This isn't exploratory R&D. It's an industrial bet: the data flywheel (more robots → more data → better model → more useful robots → more robots) only spins up if the initial dataset is massive enough to cross the threshold of general-purpose utility.


Contributor compensation: an economy of gesture

$15 million paid out to creators. The mechanics are simple: you film, you upload, human analysts validate quality and relevance, you get paid. The amount per video isn't public, but the order of magnitude suggests a few dollars per validated task — well above micro-work like Amazon Mechanical Turk.

This approach stands in stark contrast to LLM training. OpenAI, Anthropic, Google scanned the web, books, and code without paying the authors. Figure buys the data. Why? Three reasons. First: manipulation data doesn't exist on the web in sufficient quantity. Second: the required quality and diversity demand control over the collection process (sensors, instructions, validation). Third: the legal position — paying creates a clear contract, avoids lawsuits over copyright and GDPR.

The business model holds if the lifetime value of an autonomous humanoid (sales, rentals, services) far exceeds the dataset acquisition cost. At $15M for 16M raw videos, the marginal cost is negligible compared to compute budgets (H100 GPUs at ~$30k/unit, clusters of thousands of cards). The real cost isn't contributor compensation — it's the infrastructure for ingestion, curation, and continuous training.


Data infrastructure: the invisible bottleneck

Figure AI explicitly admits it: massive ingestion required an overhaul of the data infrastructure. Three major technical challenges:

24/7 availability and continuous compute. Uploads arrive continuously from 108 countries. No maintenance window. The infrastructure must scale horizontally, absorb peaks (evenings, weekends), and guarantee the integrity of heavy video files (several GB per sensor session).

Real-time feedback to contributors. A user films, uploads, and must quickly know whether their video is accepted and paid. This implies a near-real-time validation pipeline: corruption detection, format verification, first automated quality pass, then a queue for human audit. The latency of this feedback determines retention of the 44k-69k WAU.

Embedding-based deduplication. 16M videos contain massive redundancy. Filming "making your bed" produces nearly identical sequences across users. Figure uses visual embeddings to compute similarity and retain only diverse samples. The similarity threshold is a critical hyperparameter: too low → redundant dataset, wasted compute; too high → loss of useful variations (different sheets, mattresses, body types).

This infrastructure is a strategic asset. It allows Figure to iterate on Helix faster than anyone. Competitors (Tesla Optimus, 1X, Apptronik, Agility) also collect data, but few have deployed a consumer app at this scale with direct payment.


Company-reported metrics vs technical reality: reading between the lines

The 264k downloads, 16M videos, 44k+ WAU, $15M paid come from official Figure AI communications and press coverage (Humanoids Daily, Spatial Insiders). These are company-reported metrics — not audited by an independent third party. In tech journalism, they are treated as upper bounds.

Points of caution:
- WAU ≠ productive active contributors. 69k WAU (a later figure cited by Humanoids Daily) likely includes users who open the app out of curiosity without uploading a validated video.
- Videos ≠ unique data hours. 16M 30-second videos = ~133k raw hours. After deduplication, quality control, and elimination of failures (camera moved, incomplete task, object not visible), the dataset useful for Helix may be 10-20x smaller.
- $15M paid ≠ total acquisition cost. You have to add infrastructure, human analysts, training compute, engineering. The real cost per retained data hour is much higher.
- $1B commitment over 12 months: commitment ≠ actual spending. It's a budget intention, conditional on Helix's results.

This doesn't diminish the achievement. Figure deployed a complete pipeline at a rarely-seen speed: app → ingestion → curation → training → robot deployment. But rigorous analysis separates the signal (a pipeline operational at scale) from the noise (communicated round numbers).


The Competition: Tesla, 1X, Google — Who Has the Best Dataset?

Tesla Optimus leverages its vehicle fleet (cameras, FSD) and has begun collecting manipulation data in factories and labs. Advantage: massive compute infrastructure (Dojo), driving data that transfers to navigation. Drawback: fine manipulation isn't their core business.

1X Technologies (formerly Halodi) deploys humanoids in real-world environments (offices, hospitals) and collects data continuously. Their "androids as a service" approach generates operational data, not just demonstration data. But the volume remains confidential.

Google DeepMind with RT-2, RT-X, Open X-Embodiment is betting on data pooling: 22 robotics datasets, 1M+ trajectories, generalist models. It's the "Android of robotics" approach — providing the brain, not the body. Google is playing its Android of robotics: providing the brains of humanoids, not the body

Apptronik, Agility, Figure: the race for proprietary datasets. Figure has the advantage of a consumer app (Index) that scales data collection beyond the lab, into unstructured environments (real homes). That's a strong differentiator: domestic diversity > industrial cleanliness.


Paying for data creates a precedent. If Figure succeeds, others will follow. Three open questions:

Ownership of the gesture. When you film "folding a towel," who owns that movement? You (the creator), Figure (the buyer), or no one (a natural fact)? The Index contract settles it: Figure obtains a worldwide, perpetual license for model training. The contributor keeps ownership of the video but cedes the ML exploitation rights. That's clearer than web scraping, but it raises the question of the value of bodily know-how.

Collection bias. 108 countries, but what's the distribution? Probably an overrepresentation of North America / Europe, tech early-adopter users, and Western-style homes. The gestures, objects, and spatial configurations of a household in Lagos, Mumbai, or São Paulo will be underrepresented. Helix risks being less effective outside a Western context. Figure will need to actively diversify (targeted campaigns, local partnerships).

Invisible labor. The contributor wears a sensor headset, repeats tasks, waits for validation. That's work — precarious, task-based, with no social protection. $15M spread across tens of thousands of contributors = a few hundred dollars per year per active contributor. Not an income, a supplement. The legal classification (independent worker? micro-entrepreneur?) varies by jurisdiction. Figure has not communicated on labor law compliance across 108 countries.


The data flywheel in action: from the app to the deployed robot

Figure's target scheme:
1. Index collects millions of diverse human demonstrations (real homes, real objects, real noise).
2. Curation: deduplication, semi-automated annotation, selection of high-information-value trajectories.
3. Helix training: visuomotor scaling laws — more data + more compute = better zero-shot generalization.
4. Figure 02 deployment: robots in factories (BMW Spartanburg), then logistics, then homes.
5. Operational data: robots in production generate their own data (teleoperation, corrections, partial autonomy) → fed back into Helix.
6. Closed loop: better model → more useful robots → more deployments → more operational data → better model.

Index is the primer (bootstrapping). The real value will come from step 5: real operational data, not demonstration data. But without that massive primer, the model isn't good enough to be deployed usefully, and the loop never starts. This is the cold start problem of general-purpose robotics. Figure solves it by buying the cold start.


❌ Common Mistakes

Mistake 1: Confusing raw uploads with the training dataset

What's wrong: Citing "16 million videos" as the size of the Helix dataset.
The fix: Specify "16 million uploaded videos (raw data, company-reported). The effective training dataset is significantly smaller after deduplication and quality control."

Mistake 2: Equating Index with Mechanical Turk-style crowdsourcing

What's wrong: Treating contributors as interchangeable task workers, without considering the sensor + gesture + expert validation dimension.
The fix: Index collects sensorized multimodal data (video + IMU + possibly force/tactile via the headset), validated by trained analysts. This is expert curation, not unskilled microwork.

Mistake 3: Ignoring the compute cost behind data acquisition

What's wrong: Presenting the $15M paid to creators as the main cost.
The fix: Helix training compute (H100 clusters, weeks of training, continuous iterations) likely represents 10x-100x that amount. The $1B commitment over 12 months covers data + compute.

Mistake 4: Believing the data flywheel is already closed

What's wrong: Writing that Figure robots are already improving in production thanks to their own data.
The fix: At the Index stage (May 2026), the loop is not closed. Robots in testing (BMW) generate data, but the scale is not yet that of a deployed fleet. The flywheel is designed, not running at full speed.


❓ Frequently Asked Questions

How much does an Index contributor earn per validated video?

Figure does not publish the exact pricing scale. Based on $15M for ~16M raw uploads (unknown validation rate), the order of magnitude is a few dollars per accepted video. Typical monthly income remains a supplement (tens to hundreds of dollars), not a salary.

Are Index videos used only for Helix or also for other models?

Official communication: proprietary data for training Figure's models (Helix and successors). No resale, no public dataset. The monetization is internal.

What equipment is needed to contribute? Is a smartphone enough?

The Index app guides the user, but quality collection for Helix uses a sensor headset (stereo cameras, IMU, possibly tactile sensors) provided or specified by Figure. A smartphone alone does not capture the proprioception or the 3D geometry needed for visuomotor control.

Is Helix open source or accessible via an API?

No. Helix is a proprietary model, the core of Figure's competitive advantage. No open weight release has been announced. Access is via the purchase/rental of Figure 02 robots.

How does Figure handle GDPR and privacy across 108 countries?

Contributors accept terms of use that assign ML exploitation rights. Faces, license plates, and personal information visible in videos are automatically blurred before storage/training (a privacy by design pipeline). Legal compliance per jurisdiction is not publicly detailed.

The $15M (May 2026) replaces/supersedes the earlier figures. The rapid progression ($1M → $15M in a few months) reflects the acceleration of the program post-stealth.


✅ Conclusion

Figure Index is not an app, it's the industrial primer of the humanoid data flywheel. 264k downloads, 16M raw videos, $15M paid: the company-reported metrics are impressive, but the real story lies elsewhere. Figure has built an ingestion, curation, and continuous training pipeline that runs 24/7 at web scale. Helix, the visuomotor model that comes out of it, will determine whether the Figure 02 robots deliver on the promise of general autonomy. The race for proprietary datasets is on — Tesla, 1X, Google, and Apptronik each have their strategies. Figure has chosen to buy human movement at the source, in the diversity of real homes. It's a costly bet, legally cleaner, and technically ambitious. The next milestone is not the number of uploads, but the zero-shot demonstration: a Figure 02 manipulating an object it has never seen, in an environment it has never seen, without teleoperation. Homebody: Stanford pilots a humanoid with GPT-6, Astra, and a Zero Policy learned to 71%