Workers Wear Motion Trackers to Teach Robots the Jobs Humans Already Know
- Ethan Carter

- Aug 14
- 12 min read
Workers are wearing cameras and motion trackers so AI systems can study movements that took humans years to master. A recent Bloomberg account highlights a growing contest over the data behind physical work.
This effort targets a stubborn weakness in modern robotics. Language models can absorb trillions of words from digital sources. Robots cannot learn physical judgment by scraping the internet alone.
They need demonstrations that connect sight, movement, touch, force, timing, and the changing state of nearby objects. Companies are therefore turning skilled workers into live data sources for machines designed to operate in hotels, warehouses, stores, factories, and homes.
The resulting tension is difficult to miss. A worker’s expertise becomes more valuable during data collection precisely because the resulting system aims to perform more of that worker’s job.
Workers Are Becoming the Training Set
The latest robot-training pipeline begins with people performing ordinary work under extraordinary observation.
At Seoul’s Lotte Hotel, veteran worker David Park demonstrated how to fold a banquet napkin while cameras recorded his head, chest, and hands. Park had performed the task for nine years, according to an Associated Press investigation.
The exercise captured far more than the final shape of the napkin. It recorded the order of movements, how Park positioned the material, and how his hands adjusted when the fabric shifted.
South Korean robotics startup RLWRLD is converting demonstrations like these into machine-readable training data. The company also works with logistics employees at CJ and store workers at Lawson.
CJ workers demonstrate how they grip, lift, and move goods inside warehouses. Lawson employees show how they arrange products and manage food displays.
These examples illustrate why workplace data matters. A robot needs more than a video showing that a box moved from one shelf to another. It needs a usable account of how the movement occurred.
That account can include camera images, joint positions, hand poses, movement trajectories, object locations, and applied force. Each synchronized signal helps connect an observed scene to a physical action.
RLWRLD engineers add another layer by repeating tasks with virtual-reality headsets, cameras, and motion-tracking gloves. Human pilots can then guide test robots through related movements.
This process resembles imitation learning, where an AI system learns a behavior from demonstrations instead of relying entirely on hand-written rules. The principle is simple, but producing reliable examples is labor intensive.
Robots also have different bodies from their human instructors. Their joints have different limits, their hands may contain fewer fingers, and their sensors capture the world differently.
Engineers must therefore retarget a human motion to the robot’s available body. Retargeting maps the essential parts of a movement while accounting for different proportions, joints, balance constraints, and mechanical limits.
A worker might turn a wrist while folding fabric. A particular robot may need to reposition its entire arm to create a similar effect.
Even clean demonstrations do not automatically produce independent machines. Engineers must label the recordings, align data streams, remove unusable sequences, train a model, and test the resulting behavior.
The robot must then recover when an object appears in a slightly different position. That last requirement separates a memorized routine from a useful workplace skill.
RLWRLD’s own demonstrations show that gap. During one observed test, a wheeled robot moved cups near a minibar but knocked over a dish. Later footage showed a humanoid opening a box, inserting a computer mouse, and placing the closed box on a conveyor.
The progression matters, but it does not prove that the same system can reliably handle every variation of either task. It shows how demonstrations can support increasingly complex tests.
The workers are not merely labeling whether an image contains a cup. They are supplying sequences that reveal how hands, objects, and environments interact over time.
That makes human movement a new category of AI training material. It also turns workplaces into potential data-collection sites.
Physical AI Has a Data Problem the Internet Cannot Solve
Robot hardware is advancing faster than the supply of grounded, diverse, and legally usable demonstrations.
Modern AI development benefited from an enormous store of text, images, audio, and video already available online. Physical AI does not have an equivalent record of actions tied to robot controls.
An online cooking video may show someone chopping an onion. It usually lacks exact wrist orientation, applied pressure, blade position, and synchronized data describing how the onion changed after each cut.
A robot policy needs relationships between observations and actions. A policy is the learned component that selects a movement after interpreting the current situation.
Researchers can collect those relationships by operating robots directly. However, robot teleoperation requires equipment, trained operators, controlled environments, and time on expensive hardware.
Human demonstrations offer another route. Workers already possess the timing and judgment required for the task, and the workplace provides realistic objects and conditions.
The challenge is translating their behavior into signals a machine can use. Cameras capture visible movement, but they may miss force, contact, or motions hidden behind an object.
Motion trackers estimate the position and orientation of body parts. Sensor gloves can add finger movement, while force sensors measure pressure or contact.
No single device captures the complete essence of physical skill. Engineers combine signals because a visually similar motion can produce a very different result when force or timing changes.
Consider placing a glass on a shelf. The camera records the arm’s path, but safe placement depends on grip strength, deceleration, shelf height, and the moment of release.
A human adjusts these variables almost unconsciously. A robot must infer or measure them while staying within its mechanical limits.
Google DeepMind’s Open X-Embodiment project demonstrated another part of the problem. Researchers combined data from 22 different robot types through collaboration with 33 academic laboratories.
The resulting robotics dataset included more than 500 skills and 150,000 tasks. Its purpose was to help models transfer knowledge across different robotic bodies and environments.
That project used robot-generated experience rather than simply recording workers. Still, it confirms the industry’s central challenge: useful robot behavior depends on diverse, structured physical data.
A model trained on one arm, one room, and one set of objects often struggles elsewhere. The same action can look different when the camera angle, lighting, tool, or robot body changes.
Diversity therefore matters alongside volume. Developers need examples from different workers, locations, object arrangements, and task outcomes.
They also need failed demonstrations. A dataset containing only perfect movements may not teach a robot how to detect trouble or recover safely.
This requirement makes workplace collection attractive. Real environments contain wrinkles, clutter, interruptions, damaged packaging, unusual objects, and countless other conditions that laboratories struggle to recreate.
Yet realism increases privacy and quality risks. A head-mounted camera can record colleagues, conversations, documents, computer screens, customers, and restricted areas.
Data collectors must decide what enters the training set. They also need reliable records showing who consented, what was captured, and how the footage can be reused.
The process resembles the early growth of data annotation for language and vision models. The difference is that physical demonstrations capture practical expertise alongside personal and environmental information.
A recorded movement can reveal how a particular worker solves a problem. Repeated footage can expose productivity patterns, physical limitations, or informal techniques that never appeared in a procedure manual.
This is why the physical AI data race is about more than adding cameras to a workplace. It is an attempt to formalize tacit knowledge, meaning expertise that people use without fully describing it.
Once formalized, that knowledge can train many machines. The value no longer remains limited to the original task or location.
Human Skill Versus Machine Scale Is the Real Contest
The core contest is not workers against robots today, but human expertise against the scale promised by reusable robot models.
A skilled employee improves through years of repeated work. Their knowledge remains attached to one person and usually benefits one workplace at a time.
A robot model offers a different economic proposition. Developers want to record examples once, refine them through training, and distribute the resulting capability across many machines.
That promise explains the intense demand for human demonstration data. A movement captured in one hotel or warehouse can become an input for a model deployed elsewhere.
Nvidia describes this strategy through its Isaac GR00T robotics platform. The company says GR00T models combine internet video, real-world teleoperation, and synthetic data.
Its GR00T-Mimic workflow turns teleoperated demonstrations into motion data for imitation learning. Another workflow can expand a small set of human examples inside simulation.
In one company-reported test, Nvidia generated 780,000 synthetic trajectories in 11 hours. The company described that volume as equivalent to 6,500 hours of human demonstration data.
Nvidia also reported a 40 percent performance improvement when synthetic trajectories were combined with real data. These figures come from Nvidia’s own motion pipeline, not an independent benchmark across the industry.
The result still explains why real demonstrations hold strategic value. Synthetic data can multiply examples, but developers first need grounded movements and realistic task definitions.
A simulation can vary object positions, camera angles, and movement paths. It cannot guarantee that the underlying behavior reflects the subtle constraints of a real workplace.
This creates a hybrid model for robot training. Workers provide authentic examples, teleoperators connect those examples to robot controls, and simulation expands the resulting dataset.
Robot developers then test the policy on physical hardware. Failures generate new data, and people intervene when the system encounters unfamiliar conditions.
Human labor therefore appears at every stage. Workers demonstrate the task, annotators clean the recordings, operators control robots, and safety teams evaluate failures.
Automation does not arrive as a machine learning alone. It emerges from a long chain of human instruction and correction.
The strongest commercial argument says this chain eventually becomes more efficient. Each new demonstration can improve a model used across multiple robots and tasks.
The skeptical response focuses on the last mile. A robot that succeeds in a demonstration can still fail when materials, lighting, or human traffic change.
Hotel work offers a useful example. Folding one napkin requires dexterity, but completing a banquet setup also involves navigation, sequencing, visual inspection, exception handling, and coordination with colleagues.
Warehouse work poses related complications. Boxes vary in weight and condition, aisles contain people, and an apparently simple grip can depend on unstable contents.
Robots need both manipulation and situational judgment. A model that handles the first requirement does not automatically satisfy the second.
This gap pressures two groups. Robot companies must show that large demonstration datasets produce dependable real-world performance, not polished trials.
Employers must decide whether data collection supports workers, redesigns jobs, or prepares selected tasks for automation. Those outcomes can coexist within one organization.
Workers face the sharpest uncertainty. Their expertise becomes an essential input, but their rights over the resulting model remain poorly defined.
Traditional employment agreements often assign work products to the employer. Motion data raises harder questions because it records the worker’s body, habits, and accumulated skill.
Is a hand trajectory merely operational data? Is a complete demonstration a form of personal data, workplace intellectual property, or both?
The answer affects consent, compensation, retention, and reuse. It also affects whether workers can challenge a dataset deployed beyond its original purpose.
Organizations already use knowledge management to preserve written procedures and institutional memory. Motion capture extends that goal into physical behavior, where the boundaries are less settled.
The resulting database can outlast the employee who supplied it. It can also separate expertise from the person who developed it.
That separation is the central reversal. The skills considered hardest to automate are becoming highly valuable because they generate the examples automation still lacks.
Better Demonstrations Do Not Remove the Labor Risks
A richer dataset can improve robot performance while increasing surveillance, consent, and bargaining concerns for the people being recorded.
A workplace camera does not see only the target motion. It can capture faces, voices, computer screens, customer information, security procedures, and moments unrelated to robot training.
A motion-tracking system can reveal more than task technique. Repeated recordings may expose fatigue, physical differences, work speed, injuries, or deviations from management’s preferred method.
This creates a purpose problem. Data collected to train a robot might later support worker evaluation, productivity monitoring, safety investigations, or other management decisions.
Meaningful consent requires more than asking an employee to wear a device. Workers need to know what the system records, who receives the data, and how long it remains available.
They also need to know whether refusal affects scheduling, promotion, or continued employment. Consent becomes questionable when a worker cannot decline without risking income.
A 2026 study of worker-centered warehouse robots found broad interest in useful automation but also documented privacy and consent concerns. Participants wanted robots to recognize when people needed personal space.
The warehouse study also emphasized worker involvement in design. Its findings suggest that technical acceptance depends partly on control, transparency, and workplace context.
Robot-training programs need similar safeguards. Data collection should have a defined purpose, limited access, documented retention, and a process for removing unrelated material.
Bystanders present another challenge. Customers and colleagues may appear in first-person video even when they never agreed to join a training dataset.
Blurring faces can reduce some exposure, but it does not solve every problem. Uniforms, voices, locations, and work routines can identify people indirectly.
Motion data itself can also be distinctive. A dataset may reveal physical characteristics or behavioral patterns even after names are removed.
Companies must therefore treat provenance as a core feature. Provenance records where data came from, what permissions apply, and how each asset changed during processing.
Without provenance, a robot developer may receive useful files without knowing whether the people and locations were properly cleared.
Compensation presents a separate dispute. A worker might receive regular wages while demonstrating a task, a one-time payment, or no additional payment.
The dataset can then support a model with value far beyond the original shift. Existing wage arrangements rarely account for repeated commercial reuse of captured expertise.
This does not mean every worker must receive royalties from every robot action. It does mean employers and developers need an explicit position before collection begins.
Collective bargaining can influence that position. Worker representatives can negotiate device limits, payment, access rules, safety testing, and restrictions on using recordings for discipline.
The Korean Confederation of Trade Unions has warned that broad robot deployment could interrupt the development of future skilled workers. Its concern reaches beyond immediate job losses.
Junior employees learn by observing experienced colleagues and practicing under supervision. If automation removes entry-level tasks, fewer people may advance into roles where deeper expertise develops.
A training dataset can preserve an expert’s current method. It does not guarantee that the system will generate new craft knowledge when materials, tools, or customer expectations change.
Humans continually modify techniques. A static training corpus can freeze one version of a job while the real work continues evolving.
RLWRLD has argued that human mastery remains important even when AI replicates existing abilities. That distinction deserves attention.
Replication and invention are not identical. A robot may copy a demonstrated sequence without understanding why an expert departed from standard procedure.
Safety adds another limit. A robot trained on average behavior may encounter a rare situation where the expected motion becomes dangerous.
Developers need evaluation methods that cover unusual objects, nearby people, mechanical wear, sensor failures, and adversarial conditions. A successful demonstration offers little evidence about those extremes.
Workers should participate in those evaluations because they recognize workplace hazards that benchmark designers may overlook. Their contribution should not end when the cameras stop recording.
A worker-centered approach would treat employees as domain experts throughout design, testing, and deployment. It would not reduce them to anonymous sources of movement.
What Comes After the Motion-Capture Shift
The next phase will be measured by robot reliability, documented data rights, and evidence that models transfer beyond controlled demonstrations.
The first signal to watch is performance in active workplaces. Carefully edited demonstrations and laboratory trials cannot establish whether a system works across full shifts.
Developers need to report completion rates, human interventions, recovery performance, and safety incidents. Results should cover changing objects and realistic interruptions.
Independent evaluation would strengthen those claims. Company demonstrations remain useful, but they naturally show tasks and conditions selected by the developer.
A robot that folds one napkin under supervision is not yet a general hotel worker. A system that closes one box is not yet an autonomous warehouse operation.
The second signal is transfer across bodies and locations. A reusable foundation model should apply lessons from one robot or site to another with limited retraining.
DeepMind’s cross-robot research and Nvidia’s mixed-data strategy both aim at this problem. Progress requires more than scaling a single collection program.
Developers must show that training on one set of workers improves unfamiliar tasks, objects, or environments. Weak transfer would reduce the economic advantage of gathering massive demonstration libraries.
Strong transfer would make data ownership more consequential. A single recorded workflow could influence machines far beyond the workplace where it originated.
The third signal is the appearance of enforceable data rules. Companies increasingly describe consent, privacy, and provenance as features of their collection pipelines.
The important test is whether those commitments become contract terms, audit records, deletion processes, and worker rights. Public assurances alone do not settle how data gets reused.
Regulators may also examine whether recordings qualify as biometric or personal information. The answer can differ across jurisdictions and capture methods.
An ordinary video, a detailed hand model, and an ultrasound image of tendons do not expose identical information. Employers should not treat them as interchangeable.
Technical advances will complicate the distinction. MIT researchers have developed an ultrasound wristband that images muscles, tendons, and ligaments while a person moves their hand.
An AI model translates those images into positions for the palm and five fingers. The team tested the system with eight volunteers and tracked 22 degrees of freedom.
The researchers demonstrated control of a robotic hand playing a simple piano tune and operating a desktop basketball game. They are gathering data from more hand shapes and sizes.
MIT professor Xuanhe Zhao said the approach could supply large volumes of training data for dexterous humanoid robots. The ultrasound wristband also shows how intimate future motion datasets may become.
A head camera records what the worker sees. An ultrasound device records internal movement beneath the skin.
That technical progression makes clear policies more urgent. The more precisely a device captures human motion, the more valuable and potentially sensitive its output becomes.
The coming months should reveal whether physical AI companies can turn better demonstrations into repeatable operations. Watch for deployments reporting sustained performance rather than short trials.
Also watch labor agreements around wearable capture. Detailed provisions on consent, secondary use, retention, and compensation would indicate that worker data rights are becoming operational requirements.
Finally, watch whether robot models generalize across sites and machines. If they do, human demonstrations will become a strategic asset comparable to major AI training corpora.
If they do not, physical AI will remain dependent on costly site-specific engineering and continuous human intervention.
Workers are teaching machines more than isolated movements. They are exposing the unwritten structure of physical work, one grip, turn, correction, and recovery at a time.
The key question is no longer whether that knowledge can be recorded. It is whether organizations can use it without erasing the people who created it.
Developers, employers, and workers should demand the same evidence before accepting broad claims: reliable operation, transparent data provenance, and meaningful control over captured expertise.
That standard would not stop robot learning. It would make the physical AI economy account for its most important source of intelligence, the human beings already doing the work.


