Why these four thresholds.
Every supplier says its data is high quality. Almost none publish the number that decides it. Here are ours, with the derivation — including the one that is a convention rather than a result.
Four gates run at ingest. A clip that fails any of them is rejected or flagged; nothing that fails is quietly mixed into a delivery. The thresholds are not chosen for how they sound. Three follow from the optics and timing of head-mounted capture, and one does not.
1 · Hand visibility, 50%
Below half the frames, a hand-pose track contains more interpolated frames than observed ones. MANO fitting on an occluded hand is extrapolation from the last good observation, and per-joint contact state inherited from an extrapolated pose is a model output, not a measurement.
Fifty per cent is where the balance tips. It is a floor, not a target: clips that clear ingest typically run 70 to 90 per cent, and the sample records we publish sit in that band. What the floor protects is the meaning of the L5 contact layer — if contact labels can be produced on clips where the hands are mostly unseen, the label stops being evidence of anything.
2 · Inertial drift, 1 ms per 60 s
At 30 fps a frame occupies 33.3 ms. One millisecond is three per cent of a frame. The threshold is expressed per minute because drift accumulates: a three-minute clip at the limit has drifted 3 ms, still under a tenth of a frame; a nine-minute clip has drifted 9 ms, about a quarter.
The reason the budget is a fraction of a frame rather than a fraction of a second is that contact onset is a single-frame event. If the inertial stream cannot be assigned to one specific frame, it can only be assigned to a neighbourhood, and every downstream use — motion segmentation, phase boundaries, force correlation — inherits that ambiguity. Hardware timestamping is what makes the budget affordable; software synchronisation after the fact does not reach it.
3 · Reprojection error, 3 px at 1080p
This is the only gate with a direct metric consequence, so it is worth showing the arithmetic — and stating the assumption before the numbers, because at this field of view the assumption does real work.
Reprojection error is computed on the rectified image, after lens distortion has been solved and removed. A 128° raw field of view is not linear in angle per pixel; near the frame edge a naive division would be wrong by a large factor. On the rectified image, near the optical axis, the approximation holds: roughly 0.067° per pixel, so three pixels is about 0.2° of angular error. At a working distance that becomes:
Around two millimetres of lateral uncertainty at bench distance is tolerable for most manipulation and intolerable for fine assembly, where clearances run under a millimetre. So the gate is not the whole answer: fine-assembly clips carry a tighter internal target, and the published 3 px is the corpus-wide floor beneath which metric geometry stops being usable for anything.
4 · Composite score, 70 — the convention
The first three gates are necessary and not jointly sufficient. A clip can clear all three and still be unusable: the object leaves frame during the critical phase, the retarget is feasible but degenerate, the reconstruction is locally correct and globally drifting. The composite weights five sub-scores into one number: reconstruction, hand pose, tracking, retarget feasibility (does an IK solution exist within joint limits) and simulation replay (does the solved trajectory hold up under dynamics). The last two are separate because a trajectory can be kinematically solvable and dynamically impossible.
Seventy is not derived from anything. It is a working threshold, set by review rather than by measurement, at roughly the point where reviewers stopped agreeing on whether a flagged clip was usable. We would move it if that point moved, and we would say so.
We publish the distinction because a supplier that presents all four numbers as equally principled is overstating three of them or one of them. Ours is one.
Why floors rather than targets
The published work on scaling egocentric pre-training reports a log-linear relationship between hours and validation loss, with a comparatively small volume of robot data needed afterwards to align the action space. A log-linear curve has an unforgiving property: to keep improving, the corpus has to keep multiplying, and a gate that removes ten per cent of the hardest hours costs more than it returns.
That is the argument for floors. A gate should exclude data that is wrong, not data that is difficult. Every threshold above is set at the point where the measurement stops being trustworthy — not at the point where the clip stops being easy.
What the gates do not do
None of this measures whether a task is worth learning. Every gate above is a measurement-quality check on the capture. A clip can pass all four and still be a poor training example, because the manipulation is trivial, the operator’s intent is ambiguous, or the corpus already holds a thousand near-identical instances. That judgement is made separately and by people, and it is the part of data curation that has not been automated in any pipeline we know of, including ours.
There is also a failure mode in the other direction, which is worth naming because the incentive runs towards it. Tightening these gates would raise every published number and produce a corpus of clean, well-lit, unoccluded, slow manipulations — which is to say, a corpus with the hard cases removed. Occlusion, fast motion and clutter are not defects in egocentric data. They are the conditions a deployed policy will meet.
Questions this raises
A quality claim that cannot be checked is not a quality claim. Publishing the number, the unit and the derivation lets a buyer verify the figure against a delivered clip instead of taking the supplier’s word for it.
No. There is no published standard for egocentric manipulation data at the time of writing. Three of these four thresholds are derived from the physics of head-mounted capture; the fourth is a convention, and we say which.
It is rejected at ingest or flagged, and never silently mixed into a delivery. The rejection rate by scene family is reported on the data card that ships with every delivery.
No, and this is the common mistake. A gate set too tight removes occlusion, fast motion and clutter — the cases a manipulation policy most needs to see. Each of these four is a floor for usability, not a target to maximise.
Whether the task is worth learning. Every gate here is a measurement-quality check. A clip can pass all four and still be a poor training example because the task is trivial, the intent is ambiguous, or the same manipulation already appears a thousand times in the corpus. That judgement is made separately, by people.