MV-RSCH-011Research note · September 2026

Human2Robot: the embodiment gap is partly a representation problem

A human hand trajectory is not a robot action. What transfers is contact, object motion, constraints and outcome.

Human-to-robot transfer is not about copying human motion. It is about extracting the interaction invariants that survive a change of embodiment.

Figure 1 What transfers is contact, object motion and constraint satisfaction — not the hand trajectory itself.
Human hand trajectory mapped to robot action through contact, object motion and constraints

Why direct imitation of human trajectories fails

Human hands have particular degrees of freedom, finger structure, wrist range and tactile capability. The robot may be a parallel gripper, a three-fingered hand, a five-fingered dexterous hand or a dual-arm system. Mapping human joint angles onto robot joints tends to produce unreachable configurations, collisions or unstable grasps.

The failure is not in the accuracy of the mapping. It is that the wrong quantity is being mapped.

Looking for the invariants

The transferable variables are the target object, the contact point and normal, the intended object trajectory, support relations, tool constraints and the goal state. They describe task physics rather than human anatomy.

A transfer pipeline should recover invariants first and only then solve inverse kinematics, grasp and trajectory for a specific morphology. Reversing that order carries human motion style across as though it were a task requirement.

Three lines of 2026 evidence

The first is kinematic. EgoScale found that a policy trained on a 22-DoF dexterous hand transferred effectively to lower-DoF hands, supporting the view that large-scale human motion provides a reusable, largely embodiment-agnostic motor prior.

The second is visual. Human-to-robot augmentation pipelines estimate hand pose, retarget motion to simulated arms, remove human limbs by segmentation and inpainting, and composite rendered robot embodiments back into the original frames. Reported gains run 1.3% to 10.2% in simulation and 3.3% to 23.3% on real platforms. The embodiment gap, in other words, is partly a rendering problem — and rendering can only fix what the geometry recovery got right.

The third is behavioural. A model announced in August 2026 takes a single egocentric human video as its task prompt and executes previously unseen long-horizon tasks without weight updates, reported at 66% success against 9% for a language-conditioned equivalent. The result is vendor-reported, with no public weights or paper and a self-ablation baseline, so it should be read as a direction rather than a benchmark.

If that direction holds, the role of human video changes categorically. It stops being only pre-training corpus and becomes a task specification at inference time — which implies a demonstration whose viewpoint, hand visibility and task boundaries are clean enough to be read as an instruction. That is a different quality standard from bulk corpus, and nobody has defined it yet.

Retargeting is not the end. Physics validation is.

A trajectory that is geometrically reachable may still be physically unexecutable. Collision, contact stability, joint limits, object dynamics and task completion have to be checked in simulation or a digital twin.

Failures should be written back as training signal, separated into three causes: missing representation, retargeting solver failure, or an embodiment genuinely unsuited to that strategy. Each implies a different remedy, and collapsing them wastes the most diagnostic information in the pipeline.

Transferability as a third axis

Scale answers how much experience exists. Density answers how much physical information each hour carries. Transferability answers how much of that remains valid across embodiments.

It can be measured. Replay the same recovered invariants across several end effectors and report simulation pass rates. A corpus that has been through that test and one that has not are different products, whatever their hour counts say.

Three checks for your own pipeline

None of these require a vendor. All three are computable from data most teams already hold.

1. Are you transferring joint angles or contact constraints?

If the pipeline's intermediate representation is human kinematics, morphology changes will keep breaking it. If it is contact and object motion, they mostly will not.

2. Do your failure logs distinguish representation failure from solver failure?

Without that split, you cannot tell whether to buy better data or write better code — and teams routinely spend a quarter finding out.

3. What is your simulation pass rate across more than one end effector?

One end effector measures a pipeline. Three measure transferability.

Robots should not copy human joints. They should inherit human interaction constraints.

References

  1. NVIDIA GEAR Lab, "EgoScale", arXiv 2602.16710, 19 February 2026 — cross-DoF transfer.
  2. "H2R: A Human-to-Robot Data Augmentation for Robot Pre-training from Videos", arXiv 2505.11920 — 1.3–10.2% simulation and 3.3–23.3% real-world gains.
  3. Skild AI, S1 announcement, 25 August 2026; NVIDIA newsroom coverage, September 2026. Vendor-reported; no public weights, API or paper.
  4. "Robot Learning from Human Videos: A Survey", arXiv 2604.27621, 2026 — on viewpoint, action and visual alignment in embodiment transfer.
FAQ

Three checks for your own pipeline

If the pipeline's intermediate representation is human kinematics, morphology changes will keep breaking it. If it is contact and object motion, they mostly will not.

Without that split, you cannot tell whether to buy better data or write better code — and teams routinely spend a quarter finding out.

One end effector measures a pipeline. Three measure transferability.

Media & research enquiries press@movas.ai Figures in this note are produced by MOVAS AI and may be reproduced with attribution.

Data that carries these layers.

An evaluation subset ships in five business days. Name the scene family and the failure mode.

Request a Subset

Data moves robots. MOVAS moves data.