Introducing Whole-Body Manipulation Data
For robots to learn useful skills in the physical world, observing hands or objects alone is not enough. A household task, an assembly process, or a long-horizon mobile interaction depends on the coordinated motion of the body, both hands, manipulated objects, and the surrounding environment. Yet the most natural and scalable source of human demonstrations—egocentric video—rarely captures the full body. Traditional motion-capture systems, meanwhile, rely on expensive equipment, controlled spaces, and specialized operators, making it difficult to cover the long tail of real-world behavior.
GenRobot.AI has built a data production system that turns egocentric observations into Whole-Body Manipulation Data. A portable, vision-only, multi-view ground-truth system first provides high-quality supervision. DFM then reconstructs structurally plausible and temporally coherent whole-body motion from Ego observations, unifying the body, both hands, objects, contact, vision, and task semantics in a common data representation. The resulting data is designed to flow directly into robot policy training and evaluation.
Under our current internal evaluation protocol, the system achieves approximately 3 cm mean whole-body error, including about 2 cm for the upper body and 3.5 cm for the lower body. More importantly, the automated pipeline moves data collection beyond fixed laboratories and into homes, commercial spaces, and industrial environments—laying the foundation for continuously producing Whole-Body Data at the 100,000-hour scale.
Why Ego Data Alone is Incomplete
Ego devices stay close to the human point of view. They are lightweight and portable, preserve natural behavior, and place few constraints on where a task can happen. A person can wear one while cleaning a room, organizing objects, moving tools, or carrying out a long sequence of actions. This makes egocentric capture one of the most promising entry points for collecting real human skills.
But the first-person view creates a fundamental challenge: the camera sees only part of the body and environment. Arms occlude one another. Bending or squatting changes the geometry between the head and torso. The lower body often leaves the field of view entirely. Fast motion and long-distance movement continuously change both camera orientation and global position. Visible information from a single frame is therefore insufficient to recover complete, stable, and physically plausible whole-body behavior.
For robot learning, these missing signals cannot simply be discarded. A policy needs to know more than what the hands touch. It must understand where the person stands, how weight shifts, when contact begins and ends, and how body motion evolves with each stage of a task. In other words, partial Ego observations must be transformed into complete Whole-Body Manipulation Data.
This leads to a more important question: what makes data truly learnable by a model? We believe model-ready Whole-Body Data must satisfy at least four conditions: complete motion structure, aligned multimodal signals, stable long-horizon coordinates, and measurable, traceable quality. Only then does a human demonstration become more than a video to watch—it becomes a data asset that can enter robot training and evaluation directly.
Results at a Glance
Based on our internal evaluations, the system has delivered results across three critical parts of the Whole-Body Data stack:
- Whole-body reconstruction. The current model achieves approximately 3 cm mean whole-body error, including about 2 cm for the upper body and 3.5 cm for the lower body.
- Ground-truth capture. Our self-developed, portable multi-camera capture system completes on-site calibration in under 30 seconds, with calibration reprojection error below 0.4 pixels.
- Production at scale. An automated pipeline connects data ingestion, reconstruction, quality validation, valid-frame filtering, and training-format export, establishing the production foundation for Whole-Body Data at the 100,000-hour scale.
Together, these results address three practical requirements for model-ready Whole-Body Data: accurate reconstruction, reliable ground-truth capture in real-world environments, and a production pipeline that can scale with growing data volume.
From Ground Truth to Scale
Our approach does not depend on a single model or a single capture device. It is an end-to-end system spanning ground-truth production, model training, and data generation at scale.

1. Building Reliable Whole-Body Ground Truth
Reliable supervision is the starting point for high-quality Whole-Body Data. We built a low-cost, markerless, vision-only ground-truth capture system that synchronizes an Ego headset with multiple portable third-person cameras observing the same behavior. Collection requires no reflective markers, motion-capture suit, or dedicated studio, allowing the system to operate in homes, offices, outdoor spaces, and other real environments.
Visual calibration boards establish a shared coordinate system between the Ego and third-person cameras. A calibration can be completed in under 30 seconds, with multi-camera reprojection error below 0.4 pixels. Cross-view pose estimation, temporal stabilization, and smoothing then produce human-mesh ground truth directly, giving the whole-body model a consistent supervision target.
2. DFM Reconstructs Complete Motion from Partial Observations
At production scale, DFM converts Ego video into continuous whole-body motion. It combines currently visible body evidence, motion cues from surrounding frames, and structural priors about the human body. Even when part of the body temporarily leaves the camera view, DFM continues to infer the relationship among the head, torso, hands, and lower body.
The objective is not to assemble a human-looking mesh frame by frame. It is to preserve consistent body structure, global position, and action semantics across an entire sequence. A stable global coordinate frame is particularly important: it keeps the relationships among the person, scene, and objects explicit throughout long-horizon tasks, allowing the data to represent continuous movement, contact, and manipulation.
3. More Than a Body Mesh
A robot policy needs a learnable manipulation process, not isolated human geometry. Whole-Body Data therefore goes beyond body-mesh reconstruction by aligning the following signals within the same sequence:
- whole-body mesh and global motion;
- two-hand tracking and mesh;
- objects and their spatial relationships to the body;
- the onset, duration, and release of contact;
- egocentric visual observations;
- task and action semantics.
Representing these signals in a shared spatiotemporal coordinate system and data format reduces downstream conversion and secondary fitting. The output can more directly support robot policy training, evaluation, and failure analysis.
Built for Real-World Challenges
Accuracy matters only when it remains stable during real behavior. Our internal evaluation therefore goes beyond static poses and examines the system's ability to recover human motion throughout complete action sequences.
We focus on several challenges that frequently occur in Ego-based whole-body reconstruction:
- Bending and squatting. The headset orientation and body configuration change together, requiring the system to maintain a complete relationship among the head, torso, and lower body.
- Fast motion. Motion blur and abrupt viewpoint changes make individual frames uncertain, so temporal context is needed to preserve stability.
- Long-distance and long-horizon movement. The global coordinate frame must remain consistent throughout a task, preventing relationships among the person, environment, and objects from drifting over time.
- Partial occlusion and out-of-view body parts. When arms cross, the person turns, or part of the body disappears from view, the system combines visible evidence, motion context, and human-body constraints to maintain plausible predictions.
Our self-developed ground-truth production system applies automated quality checks for bone-length consistency, keypoint discontinuities, severe occlusion, out-of-frame subjects, and mesh penetration. It automatically filters invalid frames before export, replacing large-scale manual review with dataset-wide algorithmic inspection.
An Automated Pipeline for Continuous Scale
Capturing one high-quality demonstration solves only a small part of the data problem. The harder challenge is to move thousands of real behaviors continuously into a training system under a consistent quality standard.
Our production pipeline connects data ingestion, multi-stream merging, cross-view calibration, human-mesh generation, training-format export, algorithmic quality control, valid-frame selection, and dataset analytics. Work that previously required manual review, conversion, and organization can now run end to end. The path from completed capture to training-ready output is reduced from a days-long workflow to an hours-long workflow.
Portable hardware enables parallel collection across many real environments. Automated software allows processing capacity to grow with the number of devices and amount of data, without requiring quality-control headcount to scale linearly. The system also maintains dataset versions and supports batch reprocessing: when models or quality standards improve, historical raw data can be regenerated and benefit from the new capabilities.
Together, this pipeline establishes a sustainable foundation for Whole-Body Data at the 100,000-hour scale:
From Standard Motions to Real-World Skills
A fixed motion-capture studio can provide precise data, but it cannot easily represent the complexity of the real world. Human skills take place across changing spaces, lighting conditions, backgrounds, clothing, objects, and occlusions. Only sufficiently diverse data gives robot policies a chance to learn transferable behavior rather than memorize a single laboratory setup.
Lightweight, portable Ego devices move collection from fixed spaces into homes, commercial locations, and industrial environments. From organizing, cleaning, and pick-and-place tasks to longer-horizon movement and bimanual manipulation, the same data system can continuously absorb new tasks and scenes. Real demonstrations preserve natural motion and physical plausibility, while simulation built on a common human representation can further expand lighting, background, occlusion, clothing, and appearance diversity.
What’s Next
Whole-Body Data is still at an early stage. Fast motion, severe occlusion, complex human–object contact, and longer-horizon tasks remain important areas for improvement. We will continue expanding real-world task coverage, improving robustness in difficult scenarios, and advancing the data representations and transfer methods that connect human behavior to robot policy learning.
Our goal is straightforward: a natural human demonstration should no longer remain just a video. It should become complete skill data that robots can understand, train on, and reuse.
See the Whole Body. Learn the Whole Skill.





