MV-RSCH-004Research note · August 2026

The occlusion problem: why ego video alone is not enough

The most important interactions happen where the camera cannot see them. Occlusion is a structural property of embodied data.

Ego gives natural viewpoint and scale. Multi-view and 4D give ground truth behind the occlusion. They are complements, not substitutes.

Figure 1 Where the hand hides the interaction, a single first-person view cannot recover ground truth.
Occluded contact in egocentric view, recovered through multi-view and 4D ground truth

The first-person paradox

Egocentric video is the closest available proxy for what a person sees while performing a task. But at the moment a hand actually grips an object, palm, fingers and the object's critical features occlude each other. Visual uncertainty peaks exactly as the contact event does.

A monocular model can guess. Ground truth cannot be guessed. This is a limitation of observation geometry, not of resolution or model capacity, and it does not dissolve as models improve.

Occlusion is the quiet reason ego pre-training results diverge

The 2026 comparisons of egocentric against teleoperated pre-training disagreed on the headline and agreed on the mechanism: what the pipeline manages to recover decides the outcome. Occlusion sets the ceiling on what any pipeline can recover from a single viewpoint.

This is why the same corpus can support strong world-model pre-training and weak contact modelling at the same time. The frames carry enough to learn how scenes evolve. They do not carry enough to learn what the hand was doing at the moment it mattered.

Multi-view is about observability, not aesthetics

Synchronised external cameras see what the ego view cannot; depth and calibration permit three-dimensional fusion; 4D reconstruction integrates multiple frames and viewpoints into continuous geometry.

These signals matter most for articulated objects, tool use and bimanual manipulation — which is to say, for precisely the task families that contact-rich programmes care about.

Occlusion-aware sampling puts the expensive hardware where it pays

Not every task needs twelve cameras. Occlusion risk can be estimated from ego footage first — hand-object overlap, visibility of critical object parts, contact uncertainty, motion ambiguity — and only high-risk tasks routed into multi-view capture.

That turns multi-view from a more expensive way to record into a scheduling decision, which is where its unit economics actually improve.

The strategic value of 4D is what it teaches

The largest return on multi-view capture is not permanent multi-camera recording. It is training stronger monocular and stereo recovery models. High-quality reconstructions support learning object permanence, occluded pose, depth and contact inference, which lowers the enrichment cost of every future hour.

That makes it an asset rather than a project cost: today's multi-view investment becomes tomorrow's processing capability across all monocular footage — including footage nobody has collected yet.

Three checks for your own pipeline

None of these require a vendor. All three are computable from data most teams already hold.

1. What is your hand-object overlap rate during contact windows?

Computable from any corpus with hand and object tracking. It is the most direct proxy for how much of your contact data is inference rather than observation.

2. Which of your task families are you training from a single viewpoint?

Articulated objects, tool use and bimanual coordination are where single-view recovery degrades fastest. Knowing which of your tasks fall there tells you where ground truth is worth buying.

3. Are you treating 4D capture as a cost or as a teacher?

If multi-view data is used only for the episodes it contains, most of its value is being left unrealised.

The most important part of manipulation is often the part the ego camera cannot see.

References

  1. J. Ma et al., "HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining", arXiv 2606.20521, 18 June 2026.
  2. "ActiveMimic: Egocentric Video Pretraining with Active Perception", arXiv 2606.06194, June 2026.
  3. "Robot Learning from Human Videos: A Survey", arXiv 2604.27621, 2026 — on occlusion and unobserved contact as recurring limitations of video-based learning.
  4. The curves in the figure are a structural illustration of the relationship, not measured data.
FAQ

Three checks for your own pipeline

Computable from any corpus with hand and object tracking. It is the most direct proxy for how much of your contact data is inference rather than observation.

Articulated objects, tool use and bimanual coordination are where single-view recovery degrades fastest. Knowing which of your tasks fall there tells you where ground truth is worth buying.

If multi-view data is used only for the episodes it contains, most of its value is being left unrealised.

Media & research enquiries press@movas.ai Figures in this note are produced by MOVAS AI and may be reproduced with attribution.

Data that carries these layers.

An evaluation subset ships in five business days. Name the scene family and the failure mode.

Request a Subset

Data moves robots. MOVAS moves data.