MV-RSCH-001Research note · July 2026

Contact is the hidden token of physical intelligence

Language models learn structure from tokens. One of the decisive discrete events in physical AI is contact.

Contact is not a secondary annotation. It is the variable that connects visual observation to physical causality — and inferred contact and measured force are two different things.

Figure 1 Three contact products at different cost and fidelity. Only one supports contact-regulation policies.
Vision-inferred contact, instrumented contact and measured force plotted by cost and fidelity

Why contact carries so much

Robots change the world after contact. Without it the cup is not lifted, the drawer does not open, the screw does not turn.

RGB shows the result of an action but often cannot say when contact was established, where, whether it was stable, or whether it slipped. Contact is the bridge from looking like manipulation to knowing how the manipulation worked.

From a binary flag to a contact graph

Contact versus no contact is too coarse. The useful representation is a dynamic graph over hand joint to object part, tool to object, and object to environment, overlaid with onset, duration, slip, support relations and a force proxy. For bimanual tasks it also has to express coordination and role division between the hands.

A graph degrades gracefully. One low-confidence edge does not invalidate the others, which is not true of a clip-level usable-or-not judgement.

In 2026 contact stopped being an annotation and became a sensing layer

Over the past year contact and force moved from a subsection of manipulation papers into their own research layer: generalist tactile policies that transfer across sensor types, visuo-tactile world models, force-position hybrid control inside VLA architectures, video-tactile-action models proposed explicitly as a step beyond VLAs, and open-source tactile gloves built for human-to-robot skill transfer.

An independent analysis of what video-based pre-training does and does not provide put it plainly: a representation trained on human video may encode objects and affordances without encoding the contact dynamics or force constraints that manipulation requires. That is a statement about a gap in the data, not a gap in the architecture.

Three things that should not share a name

Inferred contact events are estimates of when and where contact occurred, recoverable from video, scalable to millions of hours, and silent on force magnitude and contact normal.

Instrumented contact — gloves, synchronised multi-view, markers — costs more, gives substantially more reliable timing and location, and in some configurations yields a pressure proxy.

Measured force and tactile readings are real physical quantities. They are the most expensive and the least scalable, and they are the only signal that supports training contact-regulation behaviour.

Price, use and verifiability differ across all three. Describing them all as contact annotation is the fastest way to lose credibility with a team working on insertion or in-hand manipulation.

The boundary between inferred and measured is being attacked — with measured data

A notable 2026 development is the appearance of methods that estimate tactile and force signals from egocentric video, including bimanual tactile estimation, contact-driven learning from human demonstrations, and force estimation from surface electromyography paired with human video.

It is tempting to read this as removing the need for instrumented capture. It does the opposite. Every one of these estimators is trained and validated against measured ground truth. The better the estimation research gets, the more valuable a well-aligned corpus of real force and tactile readings becomes, because that corpus is the supervision the estimators run on.

This is the general shape of the layered strategy: inference for coverage, instrumentation for ground truth, and distillation between them. The expensive tier should not be priced on its own volume. It should be priced on how much of the cheap tier it makes usable.

Three checks for your own pipeline

None of these require a vendor. All three are computable from data most teams already hold.

1. For each contact label in your corpus, can you say whether it was inferred or measured?

If that distinction is not a field in your metadata, it is not in your training weights either, and you cannot ablate on it.

2. What is your inter-annotator agreement on contact onset?

Contact timing is where human labellers disagree most. Without an agreement number, the label is an adjective rather than a measurement.

3. Do you hold any measured force at all?

Many teams working on contact-rich tasks hold none, and are training contact behaviour entirely on visual proxies. That is a defensible choice, but it should be a choice rather than an oversight.

Contact is where pixels become physics.

References

  1. "FTP-1: A Generalist Foundation Tactile Policy Across Tactile Sensors for Contact-Rich Manipulation", arXiv 2606.13102, 2026.
  2. "OmniVTA: Visuo-Tactile World Modeling for Contact-Rich Robotic Manipulation", arXiv 2603.19201, 2026; "VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs", arXiv 2603.23481, 2026.
  3. "TouchAnything: A Dataset and Framework for Bimanual Tactile Estimation from Egocentric Video", arXiv 2605.13083, 2026; A. Adeniji et al., "Feel the Force: Contact-Driven Learning from Humans", arXiv 2506.01944, 2025; J. Yin et al., "OSMO: Open-Source Tactile Glove for Human-to-Robot Skill Transfer", IEEE RA-L, 2026.
  4. "Robots Need More than VLA and World Models", arXiv 2606.06556, 2026.
FAQ

Three checks for your own pipeline

If that distinction is not a field in your metadata, it is not in your training weights either, and you cannot ablate on it.

Contact timing is where human labellers disagree most. Without an agreement number, the label is an adjective rather than a measurement.

Many teams working on contact-rich tasks hold none, and are training contact behaviour entirely on visual proxies. That is a defensible choice, but it should be a choice rather than an oversight.

Media & research enquiries press@movas.ai Figures in this note are produced by MOVAS AI and may be reproduced with attribution.

Data that carries these layers.

An evaluation subset ships in five business days. Name the scene family and the failure mode.

Request a Subset

Data moves robots. MOVAS moves data.