Global workforce fuels humanoid robot boom—but at what cost?
The gist
Humanoid robots are learning from millions of gig workers’ videos worldwide, but the data powering this AI boom could be putting those very workers’ jobs on the chopping block.
What to know
- By early 2026, companies like Figure AI and Tutor Intelligence have amassed millions of first-person videos—often from low-wage workers in India and China—rapidly accelerating robot skill acquisition.
- Ethical concerns are mounting as gig workers, frequently unaware of the risks, supply critical training data that may ultimately automate their own jobs.
- Figure AI’s $1 billion bet on global crowdsourced data aims to launch affordable humanoid robots on subscription, raising the stakes for labor markets everywhere.
Teaching Robots by Watching Us
Humanoid robots learn complex physical skills by analyzing millions of imperfect human demonstrations, mapping our every move into robotic action plans that will take years to perfect.
Humanoid robots acquire complex physical skills primarily by observing and translating human demonstrations into robotic actions, a process that involves mapping human joint movements and sequential task steps into the robot's operational space. Ayanna Howard highlights the necessity of focusing on tactile sensing and force feedback, emphasizing that robots must master specific tasks incrementally before achieving broader general-purpose assistance. Despite these challenges, the path to fully autonomous humanoid household helpers is projected to span at least a decade due to the intricacies of sensory integration and physical skill acquisition.
By early 2026, the paradigm shifted toward leveraging imperfect but diverse human demonstration data collected from gig economy workers equipped with wearable multi-camera systems, prioritizing clear visibility of hand-object interactions over flawless task execution. This scalable approach, exemplified by Microagi’s anonymized video capture of everyday chores, enables robots to rapidly learn new skills from just a few hours of footage, allowing AI to extrapolate and generalize movements without requiring pristine or complete datasets.
Advanced data collection methodologies integrate synchronized sensory inputs and motor commands recorded during human demonstrations, employing teleoperation, motion capture, egocentric video, and synthetic data to create robust training loops. Interactive imitation learning techniques like DAgger enhance this process by incorporating human corrections during robot errors, effectively turning data production into a continuous feedback loop that refines robot performance and culminates in a sophisticated 'robot brain' capable of real-time sensor interpretation and autonomous motor control.
Recent collaborations, such as between Trossen Robotics and Stereolabs, have significantly advanced the hardware underpinning physical AI data collection by integrating high-fidelity stereo vision systems like the ZED X Mini and Nano cameras. These platforms feature factory-calibrated multi-camera arrays with vibration-resistant sensors, robust GMSL2 connectivity, and real-time NVIDIA Jetson AGX Orin processing, streamlining the capture of high-resolution, motion-blur-free visual data essential for imitation learning, reinforcement learning, and sim-to-real transfer, thereby enhancing humanoid robots’ ability to acquire and deploy physical skills in real-world environments.
Data Pipelines Drive Robot Smarts
Breakthroughs in multimodal data collection and scalable teleoperation are transforming robot learning, but economic and logistical barriers now outweigh technical ones.
The evolution of robotics data pipelines has been propelled by innovations that enhance both the scale and quality of data collection. Techniques like diffusion policies enable capturing multiple behavioral modes from the same observation while maintaining training stability, unlocking scalable imitation learning. Coupled with transformer architectures and action chunking, which predict trajectories instead of isolated actions, these advances improve motion consistency and overall robot performance, demonstrating how richer, dexterous datasets drive algorithmic breakthroughs.
Scaling data collection beyond controlled lab environments remains a pivotal challenge, addressed through economic teleoperation models and multimodal data integration. Companies like Tutor Intelligence have shown that teleoperation can be scaled profitably, while platforms such as 51World’s AperOS optimize data workflows by filtering and processing diverse inputs including video, image, and egocentric simulation data. This integration acts as a multiplier for robot learning efficiency, yet the primary bottleneck has shifted from technical feasibility to economic scalability and infrastructure logistics.
Robust robotics data pipelines increasingly rely on multimodal, synchronized data streams that encompass visual, tactile, and proprioceptive inputs, captured through advanced hardware and wearable systems. Partnerships like Trossen Robotics with Stereolabs have integrated factory-calibrated stereo vision cameras with real-time NVIDIA Jetson processing to enhance data fidelity and throughput. Meanwhile, wearable systems such as Instawork’s Instacore capture synchronized multi-camera and sensor data across diverse real-world environments like Indian warehouses and kitchens, addressing the critical need for scalable, diverse real-world skill data that remains the limiting factor in physical AI development.
Addressing the long-tail edge cases essential for reliable robot deployment demands a multi-layered data strategy combining real-world, synthetic, and human egocentric data. Simulation environments enable millions of rapid training cycles to cover rare scenarios that physical data alone cannot capture, while crowdsourced platforms like Figure AI’s Index harness global video data from over 108 countries to incorporate cultural and environmental diversity. Despite these advances, uncertainty persists about the optimal volume and mix of data needed, underscoring the importance of continuous real-world operation, failure capture, and iterative model training to close capability gaps and scale robotics data pipelines effectively.
Workers Build Their Own Replacements
Gig workers worldwide are paid to record the very tasks that will train robots to take their jobs, raising urgent questions about consent, compensation, and the future of human labor.
Human labor is indispensable in generating the diverse, high-quality training data essential for advancing humanoid robotics, with gig workers and specialized operators worldwide performing repetitive, nuanced tasks under controlled conditions. Companies like Microagi and Figure AI leverage vast networks of workers—from Indian recycling colonies to US and Mexican warehouses—equipped with cameras and sensors to capture detailed motion and force data, enabling robots to learn complex household and industrial tasks. However, despite the scale of data collection, experts emphasize that only a small fraction of raw footage is usable, underscoring the need for skilled human contributors who understand the tasks they demonstrate, such as folding shirts or cooking, to ensure data specificity and effectiveness.
The ethical landscape surrounding this human-robot training ecosystem is fraught with paradox and concern, as the very workers who generate critical data—often unaware of the broader implications—risk displacement by the robots they help develop. In India, for instance, low-income laborers earn supplemental income filming their daily routines, yet remain uninformed about how their footage contributes to AI that could render their jobs obsolete. Companies attempt to mitigate privacy risks by anonymizing videos and forbidding face capture, but questions about informed consent, fair compensation, and the socio-economic impact of automating manual labor persist, highlighting a troubling dynamic where workers effectively create the manual for their own job extinction.
The global robotics data market is rapidly expanding, with projections soaring from $11.9 billion in 2024 to nearly $50 billion by 2031, driven by platforms like Figure AI’s Index and Instawork’s Instacore system that scale real-world data collection across diverse geographies and industries. These initiatives enable millions of workers to contribute without disrupting their regular jobs, capturing environmental and cultural variability crucial for robust robot training. Yet, this growth also intensifies ethical debates around labor practices, economic redistribution, and the long-term societal consequences of replacing human workers with leased humanoid robots, as envisioned by Figure AI’s CEO Brett Adcock, who proposes affordable home leasing models that could transform labor markets while addressing care and manufacturing labor shortages.
Robots Face a Data Desert
Unlike language AI, robotics lacks vast public datasets, forcing companies to create their own high-quality, task-specific data—often with skilled human operators—to push humanoids beyond the lab.
Unlike language models that benefit from vast, publicly available datasets like the Internet, robotics AI grapples with a profound scarcity of large-scale, shared data sources. As of late 2025, experts noted the absence of an 'Internet of point clouds or camera images' for robots, forcing companies to either bootstrap their own data collection or attempt indirect inference from human videos. This fragmented landscape has spurred interest in specialized firms aiming to become the 'Scale AI for robotics,' yet the consensus is that the bulk of valuable data will increasingly originate from robots deployed in real-world settings themselves, marking a shift from human-collected to robot-generated datasets.
High-quality, task-specific data emerges as the linchpin for advancing reliable humanoid robot autonomy, with research underscoring that training on data collected from the exact robot model intended for deployment dramatically improves performance. By early 2026, it became clear that mere volume of raw data—sometimes measured in petabytes—is insufficient; instead, precision in how data is gathered, including synchronized sensor inputs and motor commands, is critical. This necessitates skilled human operators working in controlled environments equipped with sensor rigs to capture nuanced, expert-level demonstrations of everyday tasks like folding shirts, reflecting the intricate human expertise required to generate usable training data.
Hardware constraints and the complexity of real-world variability pose formidable challenges to humanoid robotics, as current systems often rely on external motion capture setups impractical outside labs and struggle with limited battery life and inconsistent object manipulation. By mid-2026, researchers highlighted the need for robots to develop onboard perception capabilities—using their own cameras and sensors—to achieve autonomous operation. Moreover, the ambition to create generalizable robot autonomy capable of adapting learned skills across diverse tasks without bespoke programming remains aspirational, with early foundation models like vision-language-action systems showing promise but still far from production-ready deployment.
Addressing the enormous tail of edge cases in industrial robotics demands a multifaceted data strategy that integrates real-world robot data, synthetic simulations, and human egocentric recordings to capture the full spectrum of operational variability. As Sergey Levine and other experts have observed, the robotics community still lacks clarity on the precise quantity and optimal mix of data needed to reach dependable autonomy, underscoring the experimental nature of current efforts. This complexity is compounded by the necessity for vertically integrated data pipelines that own the robots, tasks, and feedback loops, enabling continuous improvement tailored to each robot’s unique embodiment, since datasets and demonstrations are not transferable across different robot bodies.
India and China Fuel Data Gold Rush
Massive grassroots and government-backed efforts in India and China are capturing real-world human activity at scale, powering a global race to build the world's most capable robots.
India has rapidly emerged as a pivotal hub for embodied AI data collection, leveraging its vast workforce engaged in manual tasks that remain largely automated in wealthier nations. Initiatives like ManuData Technologies and Instawork’s Instacore wearable system are harnessing thousands of workers to capture granular first-person video data of everyday activities, effectively turning human labor into the blueprint for robot training. This grassroots, government-backed ecosystem not only fuels a burgeoning market projected to reach nearly $50 billion by 2031 but also creates novel business models that integrate local labor without disrupting existing workflows, opening new revenue streams for workers and startups alike.
China is solidifying its dominance in the embodied AI data ecosystem through a multi-pronged strategy that combines cutting-edge hardware, expansive training centers, and a robust industrial ecosystem. Companies like 51World have developed advanced platforms such as the AperEgo headset and AperOS to capture synchronized multimodal data with near-perfect accuracy, while government-backed facilities across major cities train humanoid robots via teleoperation, generating high-value datasets that can cost upwards of $148,000 for mere minutes of activity. Despite shipping 23,000 humanoid robots in early 2026, practical deployment remains limited due to data shortages, prompting innovative recruitment of 'Universal Manipulation Interface' workers who perform sensor-equipped tasks to bridge this gap.
The global race for embodied AI data is intensifying, with major players from the US, China, and India aggressively competing to amass the physical interaction datasets essential for advancing humanoid robotics from impressive demos to real-world functionality. Startups like QJ Robots and Lumos in China are pioneering new business models and platforms that translate AI capabilities into practical applications across diverse robot types, while US firms such as Tesla and Agility Robotics also vie for leadership. However, market acceptance remains cautious as robots still struggle with practical tasks, underscoring the critical importance of scalable, high-quality data ecosystems to unlock the anticipated 'ChatGPT moment' for humanoid robots, potentially arriving as early as 2028.
China’s humanoid robotics market, accounting for 97% of global shipments and 85% of demand, is currently in a developmental phase reminiscent of early-stage tech giants, focusing heavily on data generation and training rather than direct consumer sales. This phase is supported by a strategic alignment of government initiatives and industrial giants in Beijing, which fosters a competitive environment for startups to innovate in AI integration and data-driven robotics solutions. Concurrently, commercial robotics applications, particularly in retail food service automation by companies like Chagee and Nayuki, are gaining traction, signaling a maturing ecosystem where data solutions and real-world deployments increasingly intersect to drive industry growth.
Crowdsourced Data Reshapes Labor
Figure AI’s billion-dollar global data platform is training robots to handle everyday tasks, setting the stage for a seismic shift in how work is done and who does it.
By 2026, Figure AI's Index platform has revolutionized humanoid robotics training through an unprecedented crowdsourced dataset, amassing 16 million video uploads from 108 countries that capture the rich diversity of real-world environments and tasks. This global approach addresses the critical bottleneck in embodied AI foundation models, as Figure commits over $1 billion to data and compute resources, continuously ingesting more than 30 minutes of video per second from everyday people performing routine activities. Such scale and diversity in training data promise to accelerate the development of truly autonomous robots capable of operating across varied cultural and architectural contexts.
The widespread deployment of robots trained on this crowdsourced data heralds a profound transformation in the labor market, particularly by automating repetitive physical tasks that currently constrain human productivity. Figure AI envisions a future where robots are accessible through affordable subscription models—leasing humanoid helpers for $400 to $600 monthly—thereby democratizing household assistance and alleviating labor shortages in sectors like elder care and manufacturing. As Figure's CEO optimistically notes, this shift could free human workers to focus on higher-value activities, making essential services more affordable and accessible while reshaping socio-economic dynamics on a global scale.










