top of page

India Humanoid Robot Training Pays Workers Now, While Automating Their Skills

7 hours ago
15 min read

India humanoid robot training has created an unsettling bargain: workers earn extra money by recording the exact skills that automation companies want machines to reproduce.

Across recycling yards, factories, warehouses, farms, construction sites, and homes, workers are fastening iPhones or specialized cameras to their heads. The recordings capture how their hands sort plastic, stitch shoes, prepare food, handle tools, and package goods.

The work provides immediate income, sometimes to people earning modest monthly wages. Yet it also converts their experience into reusable training data that can travel far beyond the workplace where it was recorded.

That conflict sits at the center of a September 8 episode of Bloomberg’s Big Take Asia podcast. The episode builds on an August investigation into tens of thousands of workers recruited for first-person data collection across India.

The primary contest is not simply people against robots. It is workers’ immediate need for income against the long-term value extracted from their recorded skills.

Robotics companies need real-world examples because impressive demonstrations rarely translate into dependable work. A machine that walks across a stage still needs extensive training before it can sort irregular objects or manipulate unfamiliar tools.

Indian workers are helping close that gap. However, the people creating this valuable resource often have limited information about its buyers, future uses, or economic value.

India Humanoid Robot Training Starts With an iPhone

The robot-training pipeline begins with ordinary workers performing familiar tasks while a camera captures their movements from their own perspective.

Bloomberg’s worker recordings include the story of Sunita Rathore, a 43-year-old waste sorter in New Delhi. She separates plastic in Karan Vihar, a recycling district on the city’s western edge.

Rathore wears an iPhone angled toward her hands. It records her removing labels, cleaning plastic, sorting material by grade, and stacking the resulting sacks.

She reportedly earns 20,000 rupees each month from her regular work. Recording the process brings another 150 rupees for each hour filmed.

That additional income matters to Rathore. She told Bloomberg that it was helping her save for her four children’s education.

Her case also illustrates the information imbalance inside this new labor market. Bloomberg reported that Rathore did not know what the footage would ultimately support.

The assignment involves more than turning on a camera. Rathore must keep her hands visible, prevent other people from appearing, and stop recording when someone approaches or calls to her.

Similar projects use cameras mounted on a worker’s head, chest, or wrist. Some systems add motion sensors, depth cameras, tactile gloves, or inertial measurement units that record bodily movement.

The resulting material is called egocentric data, meaning video and sensor information captured from the performer’s first-person viewpoint. It approximates what a robot-mounted camera would observe during the same task.

This viewpoint matters because a conventional camera across the room can miss critical details. A worker’s body may conceal the object, the angle of a grasp, or the instant a tool touches a surface.

Head-mounted video follows the worker’s attention and hands. It reveals how objects look immediately before and after each small action.

A shoe worker does not merely “make a shoe.” The worker selects a component, positions it, changes grip, applies a tool, inspects the result, and responds to imperfections.

Data companies can divide that sequence into smaller units. Each unit may receive labels describing the objects, actions, contact points, and expected result.

That combination supports vision-language-action models. These systems connect visual observations and language instructions to physical actions, such as reaching for a package or placing it on a conveyor.

The collection effort extends far beyond recycling. Bloomberg reported activity in denim factories, fishing communities, shoe facilities, construction sites, warehouses, and other workplaces.

Workers may use an iPhone attached to a cap or headband. Other assignments require equipment designed to measure motion, pressure, or depth with greater precision.

India offers both variety and scale. Many tasks targeted by robotics developers remain labor-intensive there, even when wealthier markets already use fixed industrial automation.

The country also contains dense networks of contractors, factories, service platforms, and data-labeling businesses. Those networks make it possible to recruit workers and organize recordings across many settings.

India’s role is therefore not incidental. The country is becoming part of the humanoid robot supply chain, even though the final machines may be built or deployed elsewhere.

The valuable export is not a manufactured component. It is a structured record of human experience.

Real-World Data Has Become Robotics’ Scarce Input

Humanoid developers have learned that better hardware and larger models cannot replace detailed examples of how physical work unfolds outside a laboratory.

Language models benefited from enormous amounts of text already available online. Robots face a different problem because the internet contains fewer usable records of physical action.

A written instruction can say, “Remove the label from the bag.” It does not show how much tension to apply, where to place each finger, or how to respond when the plastic tears.

A public video may show the task from an unsuitable angle. It may also lack depth, force, timing, and object-level labels.

Robots need examples that connect perception with action. They must recognize an object, choose a movement, execute it safely, and evaluate the changed environment.

Small variations can break that sequence. Lighting changes, materials bend, packages arrive damaged, tools shift, and people enter the workspace.

That is why polished demonstrations offer incomplete evidence. A robot can succeed in a controlled room while failing when clutter, interruptions, or unfamiliar objects appear.

The egocentric data process attempts to capture those variations. First-person footage preserves object interactions and fine hand movements that distant cameras often lose.

Data companies then clean and annotate the recordings. Reviewers may identify hand poses, grasp locations, contact timing, action boundaries, and the objects involved.

A clip of a person assembling a product might become several labeled actions. Those could include selecting a screw, aligning a panel, tightening the fastener, and placing the finished item aside.

This processing turns video into a dataset that a robotics team can use. Some teams train models directly from demonstrations, while others use the footage to inform simulations or teleoperation.

Teleoperation means a human remotely controls a robot while the system records the relationship between instructions, observations, and motor commands. It produces highly relevant robot data but requires access to working machines.

Human video is easier to collect at scale. A company can send phones or camera rigs to many workplaces without supplying one robot for every contributor.

That difference helps explain the sudden commercial interest in ordinary human motion. Millions of workers can generate examples in parallel while robotics labs continue improving their machines.

Human Archive represents one part of this emerging market. The San Francisco and Bengaluru startup describes its product as human sensorimotor data for physical AI systems.

According to a May profile, the company had collected tens of thousands of hours and wanted to reach millions. It said that India was its largest collection hub within Asia.

The company reportedly raised $8.2 million and established more than 120 partnerships spanning factories, hotels, restaurants, construction sites, and delivery operations. Some partnerships had not yet become active.

Its capture equipment reportedly records downward-facing 4K video at 30 frames per second. Depth cameras and wide-angle lenses add information about hands and surrounding objects.

Human Archive says it processes the resulting footage with quality-control, hand-tracking, and reconstruction models. It also says its camera positioning rarely captures faces and that post-processing removes identities that appear.

Those statements describe safeguards, not independent proof that every recording is harmless. A downward-facing camera can still capture coworkers, documents, conversations, or proprietary processes.

The broader market includes firms such as EgoLab, Humyn AI, Objectways, FPV Labs, and other data suppliers. Their methods and labor arrangements are not necessarily identical.

What unites them is the commercial thesis. Robotics companies need more diverse demonstrations than their own laboratories can economically produce.

The industry is trying to acquire physical knowledge at internet scale. India supplies both the people who hold that knowledge and the infrastructure needed to package it.

Immediate Pay Competes With the Future Value of Work

Workers receive a one-time payment, while buyers acquire data that can be copied, combined, and reused across products and markets.

The simplest defense of this arrangement is also persuasive. A worker accepts an optional recording assignment and receives additional income for work already being performed.

For Rathore, the payment supports a concrete goal. It can improve her family’s financial position long before a capable humanoid arrives at her recycling site.

Other workers may make the same calculation. A distant risk of automation can feel less important than school fees, food, rent, or medical expenses due this month.

Yet meaningful consent requires more than saying yes to a camera. The worker must understand what is being collected, why it is valuable, and how it may be used.

That standard appears difficult to meet across fragmented supply chains. A robotics company may buy processed data from a vendor that contracted with a local aggregator.

The aggregator may work with a factory owner or labor platform rather than each contributor. Workers at the end of that chain may know little about the ultimate buyer.

The Guardian examined recording practices at six factories across five Indian states. It found workers using head-mounted cameras and smart glasses without separate compensation for creating the footage.

Some participating companies argued that factories already received compensation. They said additional payments to individual workers were unnecessary.

That position exposes the fundamental dispute. A factory provides access and scheduling, but the worker’s learned movements create the data’s informational content.

A worker may also face pressure that never appears on a consent form. Refusing a camera can feel risky when a supervisor controls shifts, evaluations, or continued employment.

The factory investigation also found uncertainty among workers about what the devices were recording. Several reportedly received no direct payment beyond occasional refreshments.

Compensation varies across projects, according to the available reporting. That variation makes broad claims about fair pay or exploitation unreliable without contract-level evidence.

Still, the underlying economic structure remains consistent. A person is paid for a recording session, while the recording can retain value long afterward.

The dataset may support one model, several models, or products sold in several countries. Copies can be licensed without asking the original worker to repeat the task.

This creates a mismatch between labor time and asset value. The worker sells hours, but the data buyer obtains something closer to reusable intellectual infrastructure.

Current labor systems rarely treat muscle memory as intellectual property. A skilled sorter, welder, cook, or garment worker usually owns neither the workflow nor a claim on automation developed from it.

Robot-training data makes that omission visible. It converts tacit knowledge, meaning expertise developed through practice, into a recorded and tradable object.

A worker might receive more money for wearing a camera than for an ordinary shift. That does not establish whether the payment reflects the dataset’s downstream value.

Residual compensation offers one possible alternative. Contributors could receive additional payments when data is relicensed or used in commercial deployments.

Another approach would fund worker training, transition support, or collective benefits through fees attached to robotics datasets. Such systems would require reliable data lineage.

Data lineage tracks where a dataset came from, who contributed to it, and how it changed. Without that record, companies cannot credibly share later value.

The near-term concern is not guaranteed unemployment. Most current humanoids remain too slow, fragile, or expensive for broad replacement of complex manual work.

The more immediate problem is bargaining power. Workers are negotiating now, before anyone knows which recordings will become commercially significant.

Buyers understand why the footage matters. Many contributors do not possess the same information.

That asymmetry allows the industry to acquire potentially valuable skills at the lowest moment in their pricing cycle. By the time a dataset proves useful, renegotiation may be impossible.

Human Demonstrations Still Do Not Guarantee Reliable Robots

More training data can improve a robot’s model, but it does not automatically solve hardware limits, safety problems, or unpredictable workplaces.

The phrase “training their replacements” captures a genuine risk, but it also compresses years of difficult engineering into a single claim.

Watching a worker is not equivalent to reproducing that worker. A robot must translate human anatomy and movement into actions its own body can perform.

Human hands contain dense combinations of joints, muscles, tendons, and touch receptors. Robotic hands have different ranges of motion, strength, sensing, and durability.

A human can infer that a container is slippery after a small movement. A robot may need tactile sensors, control software, and task-specific training to reach the same conclusion.

The problem becomes harder when the human demonstration comes only from video. Pixels do not directly reveal force, weight, friction, or the effort needed to resist an object’s movement.

Depth cameras and motion sensors fill some gaps. Tactile gloves can add pressure information, but they increase equipment costs and complicate collection.

The recorded action must also be retargeted. Retargeting converts human motion into a feasible trajectory for a robot with different dimensions and mechanical constraints.

Even a technically correct trajectory can fail in a new location. A different table height, object shape, or camera position can shift what the model observes.

These limitations explain why robotics teams combine several data sources. They use human video, teleoperated robot demonstrations, simulations, synthetic scenes, and data generated during actual robot operation.

Each method offers a different tradeoff. Human video is abundant but lacks native robot controls. Teleoperation produces precise control data but scales more slowly.

Simulation generates many examples without endangering equipment. However, a simulated object does not perfectly reproduce the material properties and disorder of a physical workplace.

Real robot operation produces the most relevant feedback. It is expensive, slow, and potentially hazardous during early development.

The industry’s need for human recordings should therefore be read carefully. It proves that human action is valuable, not that replacement is imminent.

Figure AI provides a useful comparison. The company has trained its Helix system on human demonstrations for tasks such as folding laundry and moving objects around kitchens.

Figure’s chief executive told TIME that towel folding required only 80 hours of video. Yet the company’s robot still struggled with some clothing tasks during observed demonstrations.

The home robot tests showed progress and limitation at once. A machine successfully handled some kitchen tasks but was not ready for general home deployment.

A South Korean effort offers another reality check. RLWRLD collects worker demonstrations and supplements them with virtual-reality headsets, motion-tracking gloves, and direct robot control.

During one observed demonstration, a robot carefully moved cups but knocked over a dish. The failure sounds minor until the task moves beside workers or expensive equipment.

The hotel trials also revealed a large productivity gap. A hotel room that a person cleans in about 40 minutes could take a current humanoid several hours.

Lotte Hotel nevertheless hopes robots can handle cleaning and other backstage tasks by 2029. One executive estimated that humanoids might eventually perform 30% to 40% of event-preparation workloads.

The remaining work involves human interaction and difficult exceptions. That division suggests partial automation before complete replacement.

Partial automation still changes employment. A company that automates routine steps may need fewer workers, redesign jobs, or increase expected output for each person.

It can also create new roles in robot supervision, maintenance, safety, data review, and exception handling. Those roles will not automatically go to displaced workers.

Training access, geography, language, and formal education can separate the people who lose tasks from those hired to manage machines.

Claims that humans and robots will simply collaborate therefore require scrutiny. Collaboration describes a possible operating model, not a guaranteed distribution of benefits.

The same applies to confident predictions of mass displacement. Video collection shows where companies are investing, but not when machines will become economically reliable.

The strongest conclusion is narrower. Indian workers are helping robotics companies improve their odds of automating more physical tasks.

Whether that improvement produces replacement, augmentation, or another failed pilot will depend on deployment evidence, not the size of the video archive alone.

Consent and Surveillance Are Already Present-Tense Risks

Job replacement remains uncertain, but cameras, workplace monitoring, and weak control over personal data affect workers as soon as collection begins.

A head-mounted camera does more than record a pair of hands. It can capture everyone and everything the wearer sees during a shift.

That field of view may include coworkers, customers, family members, private conversations, computer screens, production methods, and personal belongings.

The worker wearing the device might consent. People moving through the background may never know that a commercial dataset includes them.

Data companies can blur faces or discard unsuitable footage. Those safeguards operate after information has already been collected and transferred.

Collection also produces a detailed account of the worker’s performance. The video can reveal pace, pauses, errors, movement patterns, and the time required for each step.

That information is valuable for model training. It is equally suitable for employee surveillance and productivity scoring.

A factory may initially approve recording for a robotics dataset. Management could later discover that the same system offers a granular view of idle time and task completion.

The resulting pressure can arrive years before a robot. Workers might face tighter targets because cameras make every motion measurable.

This possibility is not an argument against all workplace recording. Safety reviews, skills training, and quality control can use video for legitimate purposes.

The central issue is purpose limitation. Data collected for one stated reason should not quietly become a tool for another.

A credible program should specify the recording’s purpose, buyer categories, retention period, security controls, and deletion process. Workers should know whether models trained on the data can retain derived information after deletion.

Consent should also be renewable. A worker who agrees to one project should not automatically authorize every future use of the footage.

Companies need a procedure for withdrawing before data enters a model. They must also explain when technical or contractual limits make withdrawal impossible.

Privacy claims deserve independent verification. A vendor saying that cameras point downward does not address audio, reflective surfaces, documents, or people bending into view.

Anonymization can reduce risks, but complete anonymity is difficult. A distinctive workspace, voice, tattoo, uniform, or movement pattern can identify someone without a visible face.

Security matters because raw footage may reveal commercially sensitive processes. A breach could harm both workers and participating businesses.

The industry also needs provenance standards. Robotics buyers should be able to verify that recordings came from informed contributors and lawful workplace arrangements.

Without provenance, responsible buyers cannot distinguish ethically collected data from cheaper material gathered through pressure or deception.

Independent audits could examine recruitment, consent, compensation, privacy filtering, security, and downstream licensing. Audit summaries should be available to contributors and customers.

Procurement policies may exert faster pressure than legislation. Large robotics labs can require suppliers to document worker consent and reject datasets that lack credible records.

However, voluntary standards have limits when speed and cost drive the market. Vendors that spend more on worker protections can lose contracts to less transparent rivals.

Policy discussions should therefore focus on the full data supply chain. Regulating the final robot does not address harms created while gathering demonstrations.

Worker organizations also need a place in those discussions. A consent form signed by one person cannot resolve collective changes to monitoring, workload, or employment.

The toughest question concerns who can refuse. Consent has little substance when saying no threatens a worker’s shift or relationship with a supervisor.

A better system would separate recording decisions from employment evaluations. Workers should receive clear notice, direct compensation, and protection against retaliation.

India’s robot-training boom is testing these principles before the market has settled. Decisions made now can become default practices across the global physical AI industry.

Three Signals Will Show Who Benefits Next

The next phase should be judged by deployment results, transparent labor terms, and evidence that contributors can move into better roles.

The first signal is reliable robot performance outside controlled demonstrations. Companies need to show machines completing useful shifts in busy factories, homes, warehouses, or service environments.

A meaningful test should report intervention rates, task completion time, failures, and safety incidents. A short edited video does not provide enough evidence.

If robots approach human reliability across varied settings, the replacement concern becomes more immediate. If they continue requiring constant supervision, the data rush will have outrun deployment.

The second signal is whether data suppliers publish enforceable standards for consent and compensation. General promises about privacy will not settle the issue.

Useful disclosure would identify who receives payment, how refusal works, what customers can do with footage, and whether workers share any licensing value.

Buyers should also disclose their supplier requirements. Provenance records and independent audits would indicate that labor conditions matter alongside data volume.

If the industry adopts these practices, workers may gain greater visibility and bargaining power. If secrecy remains standard, today’s information imbalance will deepen.

The third signal is whether robot deployment creates credible transition paths for the people whose work supplied the training examples.

Some companies predict that workers will supervise machines remotely. One vision imagines an experienced welder in India guiding a robotic welder operating in another country.

That outcome requires more than optimism. Workers need training, interfaces they can use, reliable connectivity, and access to the employers creating those positions.

The number and quality of such roles matter. A few well-paid supervisors cannot offset widespread task removal if each person monitors many robots.

Companies should report whether existing workers move into maintenance, oversight, quality assurance, or exception-handling jobs. They should also disclose wage and retention outcomes.

Readers should resist two simple stories. One says that humanoids will soon eliminate physical work everywhere. The other says technology always creates enough better jobs.

Neither conclusion follows from a camera mounted on a worker’s head. The recording instead reveals where economic power currently sits.

Workers own the experience, but intermediaries control its conversion into data. Robotics companies control the models and potential deployment.

That distribution can change. Contracts, procurement standards, regulation, and collective bargaining can influence who receives information and value.

Developers and enterprise buyers also have leverage. They can ask whether a dataset includes documented consent before using it in a model or product.

Teams managing their own AI projects should preserve source records, usage rights, and human decisions within a searchable AI workflow. Data governance becomes harder after materials have spread across systems.

For knowledge workers, the Indian example offers a broader warning. Automation often begins by observing how people work, long before software or machines can replace them.

The person documenting a workflow may be helping improve a tool, standardize a process, or eliminate parts of a role. Those outcomes can overlap.

India humanoid robot training is therefore not only a story about distant mechanical workers. It is a test of how society values the human knowledge inside training data.

The crucial question is not whether workers should refuse every camera. Immediate income can be rational, useful, and freely chosen.

The question is whether contributors can understand the bargain and negotiate a fairer share. Watch who controls the recordings, who gets paid again, and who receives the new jobs when robots leave the laboratory.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page