top of page

Embodied Intelligence Data Collectors Start at 200 Yuan Daily to Train Robots

Jun 30
4 min read

Updated: Jul 20

Recruiters in China are hiring part-time workers for 200 to 250 yuan a day with no degree or experience required. The role involves wearing motion-capture gear and performing repetitive physical actions so that robots can learn to manipulate objects. Companies measure applicants' height and weight during interviews to fit collection gloves and ask about VR motion sickness before assigning shifts.

Job Tasks Split Between Remote Operation and Free-Hand Recording

Two data collection methods dominate the current workflow. In remote operation sessions, workers strap on sensors and control dual-arm robots to sort blocks or stack paper cups. In free-hand sessions, workers simply repeat actions such as folding clothes while cameras and sensors record their movements for later processing. Both approaches generate trajectory data that robotic learning models require.

The process runs in controlled rooms equipped with multiple camera angles and tracking devices. Sessions last two to four hours with short breaks. Workers receive payment the same day or the next day after quality checks confirm the data meets minimum standards for smoothness and repeatability.

Global Data Shortage Drives Hiring Surge

High-quality physical interaction datasets totaled roughly 500,000 hours worldwide by early 2026. That figure represents less than one twenty-thousandth of the text data used to train large language models. The gap explains why startups and research groups accelerate manual collection rather than wait for simulation improvements alone. Bloomberg reports on data pipelines at companies including Figure AI - which uses the recordings to train its Figure 01 humanoid - and Agility Robotics, which collects similar human demonstrations for Digit platform development.

Hardware constraints compound the shortage. Current robotic platforms still need thousands of demonstrations per simple skill before achieving reliable performance. Each new object shape, surface texture, or task variation typically requires fresh recordings because simulation-to-real transfer remains inconsistent. Specific projects such as Google DeepMind’s RT-X and OpenAI’s robotics initiatives draw on these datasets for manipulation benchmarks.

Workers Face Physical and Technical Demands

Applicants quickly encounter limits once shifts begin. Holding precise postures for extended periods produces fatigue in shoulders and wrists. Some tasks require repeated fine finger movements that trigger hand cramps within an hour. VR headset users report dizziness during longer sessions, prompting questions about susceptibility during the interview stage.

Data quality rejects occur when movements deviate from expected paths or contain excessive jitter. Rejected recordings force workers to repeat sessions without extra pay. The combination of repetitive motion and strict quality gates creates turnover that recruiters offset with constant new postings.

A Shanghai-based recruiter noted, “We look for steady hand motion above all - anyone who can repeat the same grip 200 times without drift passes the first filter,” according to hiring posts reviewed by The Verge. One collector added, “After four hours of cup stacking my wrists were numb and I still had to redo two takes because of slight shakes.” Another described, “The motion-capture suit pinched during bends; I lost feeling in my fingers halfway through a clothes-folding shift.”

Data Collection Mirrors Broader Training Challenges

The same principle applies to office agents that rely on rich context to produce accurate outputs. Just as robotic models need thousands of recorded trajectories, agents need consistent records of meetings, documents, and decisions. Systems that accumulate context passively avoid forcing users to re-explain background every session.

remio captures browsing history, meeting transcripts, and file activity continuously. The resulting memory layers allow the agent to reference prior choices when generating presentations or reports. This architecture reduces the need for manual context injection in the same way large-scale robotic datasets reduce reliance on hand-crafted rules.

Quality Standards and Verification Processes

Labs apply automated checks before accepting recordings. Trajectory smoothness scores and endpoint accuracy thresholds filter out noisy sessions. Human reviewers spot-check a sample of each batch to confirm labels match the intended task. Rejected data returns to the queue for re-collection.

Payment structures reflect these gates. Base daily rates assume a target number of valid recordings. Consistent performers sometimes receive small bonuses, while repeated quality issues lead to removal from active shifts. The model mirrors gig-economy patterns where speed matters less than usable output.

Industry Implications Extend Beyond China

International robotics groups watch these collection pipelines closely because demonstration volume directly limits model scale. Projects that secure steady human recording capacity advance faster on manipulation benchmarks. Teams without equivalent access rely more heavily on synthetic data and face higher sim-to-real gaps. Reuters coverage highlights how firms such as Boston Dynamics integrate similar human-collected trajectories for Atlas robot skill training in its Waltham facility.

The labor requirement also surfaces debates about long-term sustainability. As task complexity rises, the hours needed per skill multiply. Some researchers explore whether improved imitation learning or better hardware can reduce the per-task demonstration count, yet current progress remains incremental.

Outlook for the Next Three Months

Volume metrics from major collection platforms will indicate whether daily hiring rates hold or decline. Hardware upgrades that simplify sensor calibration could lower entry barriers and widen the applicant pool. Published manipulation benchmarks that use newly released datasets will reveal whether quantity gains translate into measurable performance lifts.

Any regulatory moves around worker classification or session length limits would also alter economics for the companies running these programs. Continued growth in embodied AI funding will likely sustain demand for the foreseeable future, while plateaus in model improvement could reduce collection budgets if returns diminish.

The pattern underscores a recurring requirement across AI domains: high-quality, human-generated data remains a bottleneck even as model architectures advance. Organizations that secure reliable collection pipelines hold an advantage in training the next generation of capable systems.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page