MV-RSCH-009Research note · September 2026

Can robots learn from YouTube? A pipeline, and the arithmetic behind it

From rights gating through candidate discovery, 3D, contact and state transitions to training routing.

Open video becomes a reliable upstream for physical AI only after rights governance, filtering, structure recovery and confidence assessment — and rights governance has to come first, not last.

Figure 1 From rights gating to trainable trajectories. The arithmetic decides whether the pipeline is worth running.
Funnel from web video through rights gating, candidate discovery, 3D, contact and state transitions

Stage 0: the rights gate goes first

Putting discovery first is the wrong order. Acquisition, decoding and storage can themselves be the factual basis of a claim, independent of what is later trained. Discovering that a source is unusable after structure recovery has already run destroys the compute spend and every derived label with it.

Stage 0 covers source allowlisting and licence classification, a recorded review of terms of service, the consent status of filmed subjects and venues, and the permitted use boundary for that source — research, pre-training, or commercial derivatives. Each is written into the provenance record and travels with the segment to the end of the pipeline.

Stage 1: candidate discovery

The first stage does no expensive reconstruction. It filters cheaply for human manipulation, viewpoint, hand visibility, object salience, camera stability, task duration and scene class, isolating the small share of raw footage worth processing further.

The quality of this judgement sets every downstream cost. Compute belongs here.

Stage 2: physical parsing

Surviving candidates go through hand pose, object tracking, action segmentation, depth estimation, camera motion and object-state detection. The goal is not state of the art in every module. It is a consistent confidence output and a reliable answer to whether a sample earns the next layer.

Low-confidence segments remain useful for pure video pre-training. This is a routing decision, not a pass-fail verdict.

Stage 3: structure recovery

High-value samples get camera pose, scene geometry, object pose, contact events and transition structure.

Monocular open video means many of these quantities can only be estimates. Every derived label therefore carries provenance, method, confidence and uncertainty, and none is presented as ground truth. Research teams weight by these fields, which makes a mislabelled certainty worse than a missing label.

Stage 4: training routing

The output is not one large dataset. World-model pre-training wants scale and dynamics; VLA training wants language and action proxies; contact models want highly visible interaction; retargeting wants geometry and hand quality. Each segment goes where it is worth the most.

This is where the latent-action route fits rather than competes. Segments that survive filtering but fail structure recovery are exactly the material a latent action model can still use — provided camera motion has been characterised, which Stage 2 does anyway.

The arithmetic

End-to-end retention from raw source to structured training pool runs below one percent in our target operating model. That number has to be read together with the very high ceiling on supply.

The implication runs both ways. A million-hour raw pool yields a structured asset in the tens of thousands of hours, which means a one-point improvement in the first filter moves unit cost more than any downstream optimisation. And it explains why holding petabytes of video is a weak claim, while an auditable end-to-end usable rate is a strong one.

It is also the upstream mirror of the yield problem inside capture operations. One loses samples in manufacturing, the other in selection. Financially they are the same multiplication.

Three checks for your own pipeline

None of these require a vendor. All three are computable from data most teams already hold.

1. What is your first filter's precision and recall?

Measured on a human-reviewed set, with a sensitivity analysis showing how unit cost moves with it. In an internet mining pipeline this is the number that matters most and is measured least.

2. Are your inferred confidences calibrated against anything?

An uncalibrated confidence cannot be used for weighting, which makes it functionally equivalent to no confidence at all.

3. If a source became unusable tomorrow, could you find every derived artefact from it?

Provenance is only real if it supports recall. Most catalogues can answer where a file came from and not which downstream labels depend on it.

Don't download the internet. Mine it for robot-learnable experience — and prove where every frame came from.

References

  1. "HuRo", Conference on Robot Learning 2026 — approximately 630,000 converted episodes and 142 million processed frames from five sources; reported in Rocking Robots, September 2026.
  2. "Why Latent Actions Fail, and How to Prevent It", arXiv 2605.20223, 2026; A. Liang et al., "CLAM: Continuous Latent Action Models for Robot Learning from Unlabeled Demonstrations", arXiv 2505.04999.
  3. "ActiveMimic: Egocentric Video Pretraining with Active Perception", arXiv 2606.06194, June 2026 — on camera motion as signal.
  4. Regulation (EU) 2024/1689 (EU AI Act); actions pending in United States federal court under 17 U.S.C. §1201. No final rulings; not legal advice.
FAQ

Three checks for your own pipeline

Measured on a human-reviewed set, with a sensitivity analysis showing how unit cost moves with it. In an internet mining pipeline this is the number that matters most and is measured least.

An uncalibrated confidence cannot be used for weighting, which makes it functionally equivalent to no confidence at all.

Provenance is only real if it supports recall. Most catalogues can answer where a file came from and not which downstream labels depend on it.

Media & research enquiries press@movas.ai Figures in this note are produced by MOVAS AI and may be reproduced with attribution.

Data that carries these layers.

An evaluation subset ships in five business days. Name the scene family and the failure mode.

Request a Subset

Data moves robots. MOVAS moves data.