( 02 )Layer

4D Multi-View Ground Truth

10,000+ hours

Time-synchronised multi-camera capture on the same clip as the egocentric view.

Single-view capture loses the object the moment a hand closes around it. Six to twelve hardware-genlocked cameras surround the work volume, so every moment of a manipulation is observed from at least one unoccluded angle.

Use it as ground truth for validating whatever your own pipeline produces from monocular video — or as the supervision signal for training that pipeline.

This requires rigging and calibrating cameras on site. A contributor network capturing on personal devices cannot produce it.

SPECIFICATION
Camera rig
6–12 synchronised views
Genlock accuracy
< 1 ms hardware trigger
Calibration RMS
< 0.5 px, re-verified per session
Occlusion coverage
> 98% of frames

Available on MOVAS-Ego 100K. Coverage per scene family is published on each family page and on the data card.

FAQ

4D Multi-View Ground Truth — Common Questions

Between 6 and 12 depending on work-volume size. Dense manipulation sessions use 12; wider-area sessions use 6 to 8 with longer baselines.

Both are provided — metric depth from active sensors where the rig includes them, plus multi-view stereo depth from the calibrated array, so you can cross-check.

Your models need the physical world.

Start with training-ready data today. Scale to targeted re-capture when you’re ready.

Request a Subset

Data moves robots. MOVAS moves data.