MV-RSCH-002Research note · July 2026

The yield problem: why most physical AI data never reaches training

The value in a data programme sits in end-to-end yield, not in recorded hours.

Physical AI data is a manufacturing line. A low first-pass rate at any gate is amplified exponentially at scale.

Figure 1 Rolled throughput yield across six gates. Small per-gate losses compound fast.
Rolled throughput yield curves at 90, 95 and 99 per cent across six pipeline stages

Data is not files. It is manufactured output.

Recruitment, task design, device calibration, capture, upload, synchronisation, segmentation, annotation, reconstruction, validation and delivery — every step can lose samples.

Manufacturing has a standard name for this: rolled throughput yield, the product of first-pass rates across all operations. It multiplies rather than adds, which means local optimisation is diluted by every other stage, and broad small improvements compound far beyond intuition.

A yield tree

Six gates are enough to locate almost any loss. Capture yield covers equipment and task completion. Synchronisation yield covers multi-stream alignment. Perception yield covers hand and object trackability. Reconstruction yield covers geometric recoverability. Interaction yield covers whether contact and state are trustworthy. Robot-ready yield covers whether retargeting and simulation pass.

Six gates at 90% gives roughly 53% end to end. At 95%, roughly 74%. At 99%, roughly 94%. Moving from 95% to 99% — four points per gate — raises end-to-end yield by about 27% in relative terms. At scale that is not an operations detail. It is the cost structure.

QC has to move upstream

The most expensive defect is one discovered after capture is complete and the footage has crossed a border.

High-volume operations need lightweight checks at the edge or at the point of capture: exposure, frame rate, inertial data, synchronisation, hand visibility, task completion and storage integrity, with immediate feedback. Rework cost is non-linear, because redoing a session usually means reassembling people and a location rather than repeating a recording.

The same arithmetic governs mined data

Internet-scale mining loses material at selection rather than at manufacture, but the structure is identical: a chain of multiplicative retention rates ending in a small structured output.

This is why the two strategies converge operationally even though they look different commercially. Whether the loss happens because a camera was mispositioned or because a clip failed a visibility filter, the fix is the same — measure each gate, move the cheap gates upstream, and price on what survives.

From cost per hour to cost per valid transition

Procurement asks about cost per recorded hour. The metrics that predict anything are cost per usable hour, cost per structured hour, cost per simulation-valid episode, and eventually cost per valid state transition.

One honest qualification: cost per valid transition is a proposition, not yet an accepted method. It becomes a method when there is evidence that valid transition count relates stably to downstream model outcomes. Until then the sensible practice is to quote on comparable terms and manage internally on transition terms — and to say which one you are doing.

Three checks for your own pipeline

None of these require a vendor. All three are computable from data most teams already hold.

1. Do you know your own end-to-end yield?

Raw input to trainable asset, as a single ratio. Teams that have never computed it are usually surprised by a factor of two or more.

2. Which gate is your worst, and when did you last measure it?

In a multiplicative chain the worst gate dominates. It is also the one most likely to have drifted since it was last checked.

3. How far upstream is your earliest quality check?

If the first check happens after upload, every defect costs a re-shoot. If it happens on device, most defects cost a retake.

At scale, data quality is a yield problem.

References

  1. Rolled throughput yield is a standard manufacturing measure; its multiplicative property is arithmetic and does not depend on any particular framework.
  2. "HuRo", Conference on Robot Learning 2026 — approximately 630,000 converted episodes from five raw video sources, illustrating selection-stage retention at scale.
  3. The 90 / 95 / 99 curves in the figure are illustrative parameters, not measured yields.
FAQ

Three checks for your own pipeline

Raw input to trainable asset, as a single ratio. Teams that have never computed it are usually surprised by a factor of two or more.

In a multiplicative chain the worst gate dominates. It is also the one most likely to have drifted since it was last checked.

If the first check happens after upload, every defect costs a re-shoot. If it happens on device, most defects cost a retake.

Media & research enquiries press@movas.ai Figures in this note are produced by MOVAS AI and may be reproduced with attribution.

Data that carries these layers.

An evaluation subset ships in five business days. Name the scene family and the failure mode.

Request a Subset

Data moves robots. MOVAS moves data.