( 01 )Layer

L1–L5 Control Annotation

Full corpus

The supervision a policy actually trains on — not the labels a video search engine needs.

Most video annotation answers what a clip contains. Control annotation answers what a policy has to reproduce: which phase the operator is in, which object is involved, which joints are touching it, and whether the attempt succeeded.

A first engagement typically starts at L1–L3 and adds L4–L5 once training value is demonstrated.

ANNOTATION DEPTH LADDER
L1   Task
Task identity and success/failure outcome.
L2   Phase
approach → grasp → transport → manipulate → release.
L3   Object
Open-vocabulary object identity, material and state transitions.
L4   Keyframe
1 fps dense keyframes plus every phase-boundary frame.
L5   Contact
Per-hand, per-joint contact points and duration, derived from MANO surface distance.
SPECIFICATION
Coverage
L1–L3 across the full corpus
Review
Expert adjudication on a sampled subset
Agreement
Published per delivery on the data card

Available on MOVAS-Ego 100K. Coverage per scene family is published on each family page and on the data card.

FAQ

L1–L5 Control Annotation — Common Questions

Video-understanding labels are optimised for retrieval and captioning. They describe a clip; they do not tell a policy which phase it is in or which joints are in contact. Those are the two signals that determine whether a demonstration is learnable.

From MANO surface distance against the reconstructed object mesh, then adjudicated by expert reviewers on a sampled subset. Precision and recall against that sample are published on the data card.

Your models need the physical world.

Start with training-ready data today. Scale to targeted re-capture when you’re ready.

Request a Subset

Data moves robots. MOVAS moves data.