Robot-ready is not a label. It is a measurable readiness profile.
Conventional data QA proves a media file is sound. It says nothing about whether the data can train a robot.
Robot-ready is not binary. It is a profile spanning visual usability, physical structure, coverage, novelty and transferability.
Why conventional QA is insufficient
Sharp footage with no dropped frames proves the media file is sound. It proves nothing about training value.
Robot data QA has to answer six questions at once: is the critical interaction visible, is the state change parseable, is the object trackable, is contact determinable, does this add new coverage, and can it be validated through retargeting or simulation. Conventional video QA answers none of them.
A profile, not a black box
Eight interpretable sub-metrics are enough to start: visual integrity, interaction observability, geometry recoverability, state completeness, contact confidence, novelty, transfer feasibility, and rights and provenance. Different training objectives weight them differently — a world-model pre-training run and a contact-policy fine-tune want almost opposite profiles from the same catalogue.
There is a practical reason to show a profile rather than a total. A single score invites comparison against other single scores, which is a conversation with no information in it. A profile forces the question of what is being trained.
Downstream calibration is the only source of credibility
A readiness score is credible because it correlates with something, not because of what it is called. The test is whether high-scoring data produces lower validation loss, higher task success, or a reduced requirement for fine-tuning robot data. Any sub-metric with no relationship to downstream performance should be down-weighted or removed.
The form of evidence that carries weight is specific: a held-out evaluation protocol that can be reviewed, an ablation curve with and without the batch, and — most persuasively — an exchange rate showing how many hours of teleoperated data the batch substitutes for.
The field already has partial versions of this. The 2026 comparisons of egocentric and teleoperated pre-training are exactly exchange-rate experiments, run on pre-training sources rather than on individual batches. Extending that method from data type to data batch is the open problem, and it is more valuable than any scoring rubric.
What a delivery report should contain
Coverage against a declared taxonomy. Quality and confidence distributions per layer. End-to-end yield for the batch. Provenance. And known limitations, stated rather than omitted.
The last item is the one most often missing and the one most predictive of whether the rest is trustworthy.
Standards come after evidence, not before
A measurement framework published before its correlation evidence is marketing wearing a lab coat, and technical readers identify it immediately.
The sequence that works is internal reporting first, correlation experiments second, published methodology third. We are at the first step and say so.
Three checks for your own pipeline
None of these require a vendor. All three are computable from data most teams already hold.
1. Can you predict, before training, which batch will help?
If not, every data purchase is an uncontrolled experiment. The goal of a readiness profile is simply to make that prediction possible and then check it.
2. What is your exchange rate between data types?
How many hours of human video substitute for one hour of teleoperation, on your tasks and your platform. Published comparisons give a starting prior; only your own ablation gives a number.
3. Does your quality report include known limitations?
A report with no limitations section has not been written by anyone who ran the pipeline.
Robot-ready is not a label. It is a measurable readiness profile.
References
- J. Ma et al., "HumanScale", arXiv 2606.20521, 18 June 2026 — matched-scale comparison of pre-training sources.
- "ActiveMimic: Egocentric Video Pretraining with Active Perception", arXiv 2606.06194, June 2026 — opposite headline result under a different pipeline.
- NVIDIA GEAR Lab, "EgoScale", arXiv 2602.16710, 19 February 2026 — validation loss correlating with downstream real-robot performance.
- The three profiles in the figure are constructed illustrations, not delivered data.
Three checks for your own pipeline
If not, every data purchase is an uncontrolled experiment. The goal of a readiness profile is simply to make that prediction possible and then check it.
How many hours of human video substitute for one hour of teleoperation, on your tasks and your platform. Published comparisons give a starting prior; only your own ablation gives a number.
A report with no limitations section has not been written by anyone who ran the pipeline.
Five supply models, not one market
Why raw hours went to zero, how supply split into five structurally different businesses, and what became scarce instead.
ReadHuman2Robot: the embodiment gap is partly a representation problem
A human hand trajectory is not a robot action. What transfers is contact, object motion, constraints and outcome.
Read