MV-RSCH-006Research note · August 2026

The physical data scaling law: what to scale, and where the curve bends

Scale still matters. What needs scaling is the state space, not the hour count.

Hours are a supply metric, not a learning-value metric. The next data infrastructure has to optimise scale, coverage and marginal information gain at once.

Figure 1 Three published scaling results side by side. They describe different regimes of one curve.
EgoScale, a VLA foundation model and Dyna-2 scaling results compared across pre-training hours

Three independent curves, and they do not say the same thing

NVIDIA's GEAR Lab published EgoScale in February 2026: a vision-language-action model trained on 20,854 hours of action-labelled egocentric human video, with a log-linear relationship between human data scale and validation loss, and that loss correlating strongly with real-robot performance.

A VLA foundation model reported in 2026 scaled pre-training from 3,000 to 20,000 hours and observed both success rate and progress score improving throughout, with no sign of saturation at the top of that range.

Dyna-2 went two orders of magnitude further and found the trend continuing — but with a small fitted exponent, and with most of the visible gain arriving between ten thousand and one hundred thousand hours.

Put side by side these are not contradictory. They describe different regimes of the same curve. Below roughly ten thousand hours, returns are steep and nothing is saturating. Around a million, the curve is still rising and it is flat enough that the next order of magnitude is an expensive way to buy capability.

What that means at the purchase order

Convert the Dyna-2 ladder into marginal terms: 1K to 10K buys eight points, 10K to 100K buys seventeen, 100K to 1M buys eight. For a team that already holds a hundred thousand hours from a similar distribution, the expected return on another nine hundred thousand similar hours is about what the first nine thousand delivered.

For a team holding five thousand hours, the opposite advice applies: buy more, almost regardless of what. The useful question is not whether to scale. It is which regime you are in, and — once past the steep part — what you are scaling into.

Diversity as a set of auditable dimensions, not a score

A physical diversity index is only useful if it decomposes. Scene, object, state, interaction, embodiment, and failure-recovery diversity each answer a different version of the same question: which gap did this hour fill?

A single composite number invites comparison against other single composite numbers, which is a conversation with no information in it. A decomposition invites the question of what the team is trying to train, which is the conversation worth having.

We use such an index internally for scheduling. We are deliberately not publishing it as a standard, because until there is evidence that it correlates with downstream training outcomes, publishing it would be marketing rather than measurement.

Marginal information gain sets the price of hour N+1

If a corpus already contains two hundred thousand ordinary drawer-opening episodes, one more adds almost nothing. A jammed drawer, an off-axis grasp, a one-handed attempt because the other hand is occupied, a recovery after the object slips — each of these may fill a genuinely new interaction state.

A useful scheduler estimates that gain before collection, from corpus embeddings, task ontology, state taxonomy and quality signals, and allocates capacity accordingly. This is the difference between running a collection programme on labour cost and running one on judgement.

There is a useful adjacent result from language modelling: repeated data holds roughly the value of fresh data for about four epochs and decays quickly after that. No equivalent threshold has been published for physical data, but the structural point — repetition has a ceiling — extrapolates reasonably.

From buying hours to buying capability coverage

What a model team ultimately needs is capability, not footage. The useful specification is how many manipulation families, state transitions, contact modes and long-horizon sequences are covered, and how that coverage maps onto training objectives and known failure clusters.

This is also how pricing power moves. Priced by the hour, data converges on the marginal labour cost of collection, and that only goes one direction. Priced by coverage, it is anchored instead to what closing a specific gap would cost the buyer to do themselves.

Three checks for your own pipeline

None of these require a vendor. All three are computable from data most teams already hold.

1. Which regime are you in?

Below about ten thousand hours of in-distribution data, scale is still the cheapest capability you can buy. Above a hundred thousand, it usually is not. Knowing which side you sit on changes the entire procurement question.

2. Can you measure what the last purchase added?

Not whether the model improved — whether the corpus's coverage changed. If a new batch does not move any coverage dimension, it will not move the model either, and you will find that out much later.

3. What is your corpus's redundancy rate?

Very few teams measure how many near-duplicate episodes they hold. It is usually the single highest-leverage number in a data budget, and it is computable from embeddings you already have.

The next scaling law is not hours alone. It is the expansion of physical state space.

References

  1. NVIDIA GEAR Lab, "EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data", arXiv 2602.16710, 19 February 2026; project page, research.nvidia.com/labs/gear/egoscale; publication listing, UT Austin Robot Perception and Learning Lab.
  2. "A Pragmatic VLA Foundation Model", arXiv 2601.18692, 2026 — pre-training scaled from 3,000 to 20,000 hours with no observed saturation.
  3. Dyna Robotics, "Dyna-2: A 1-Million-Hour Scaling Law for World-Action Models", research page, 15 August 2026; independent technical write-up, MarkTechPost, 13 August 2026.
  4. N. Muennighoff et al., "Scaling Data-Constrained Language Models", NeurIPS 2023 — on the value of repeated data.
FAQ

Three checks for your own pipeline

Below about ten thousand hours of in-distribution data, scale is still the cheapest capability you can buy. Above a hundred thousand, it usually is not. Knowing which side you sit on changes the entire procurement question.

Not whether the model improved — whether the corpus's coverage changed. If a new batch does not move any coverage dimension, it will not move the model either, and you will find that out much later.

Very few teams measure how many near-duplicate episodes they hold. It is usually the single highest-leverage number in a data budget, and it is computable from embeddings you already have.

Media & research enquiries press@movas.ai Figures in this note are produced by MOVAS AI and may be reproduced with attribution.

Data that carries these layers.

An evaluation subset ships in five business days. Name the scene family and the failure mode.

Request a Subset

Data moves robots. MOVAS moves data.