Five supply models, not one market.
In eighteen months the price of a raw egocentric hour went to zero and the supply side split into five structurally different businesses. Most provider comparisons still print them in one table as though they were substitutes.
The useful question is no longer which supplier is largest. It is which of five supply models a buyer is actually contracting with, because they differ in what can be licensed, how fast it arrives, and whether it can be used commercially at all.
1 · What changed
Through 2025 the category behaved like an ordinary data market: more hours cost more money, and the largest corpus won. That ended when permissively licensed supply reached the size of a large commercial corpus.
Build AI’s Egocentric-100K carries 100,405 hours under Apache 2.0, free for commercial use, verifiable on the dataset card. Its 10,000-hour predecessor draws roughly 140,000 downloads a month. Over the same period, published teleoperation capture rates fell by around two-thirds.
Supply indexed to the largest permissively licensed release available each quarter. Price indexed to published per-hour teleoperation capture rates over the same period. Both normalised to their own maximum; the shapes matter, the units do not.
What is free is not equivalent to what is sold. The 100K release ships at 256p with no inertial stream, no hand pose and no contact labels; the 10K release is 1080p but an order of magnitude smaller. The point is narrower and harder to argue with: a paid hour now has to be different from a free hour, not merely more numerous.
2 · Five supply models
Underneath the vendor names there are five ways of obtaining an egocentric hour, and they produce legally and technically different objects.
Open and academic corpora
Sets the floor price at zero. Academic corpora cannot be used commercially at all, which is the most common misreading in the category.
Gig workforce networks
Wins on unit cost by an order of magnitude. Rights stop at the individual, which is sufficient for public and domestic settings and insufficient for workplaces.
Contributor networks with deep annotation
Currently the deepest published annotation and the largest accumulated volume. Rights structures are generally not published.
Platform and OEM in-house
The largest single source of demand leaving the merchant market. What is collected serves one buyer and is often bound to a platform or to state capital.
Employer-signed third-party capture
The only model in which a buyer can be granted rights over premises, processes and the colleagues in frame. Slowest to scale, because every site is a negotiation.
These are not tiers of quality. A gig network is the right supplier for public-space navigation data and the wrong one for a commercial kitchen. An open corpus is the right starting point for a visual backbone and cannot be the basis of a deployment programme. Printing them in one ranked list obscures the only distinction that decides a purchase.
3 · Demand and capacity are in different places
The category is often described as consolidating around China. That is true of capacity and untrue of revenue, and conflating the two produces bad forecasts.
Booked revenue and the named frontier buyers concentrate in North America, where the specialist market analyses put the largest regional share. Operated collection capacity concentrates in Asia, with the largest single-operator throughput figures published by Chinese platforms and the lowest unit cost published by Indian operators. New capital in the first half of 2026 concentrated heavily in China.
Figures are as published by operators or in named coverage and are not directly comparable: the China figure is a machine-side rate rather than a wage, and coverage, equipment and QA differ by market. The order of magnitude is the point.
The structural consequence is straightforward. A supplier whose cost base and revenue base sit in the same high-cost region is competing against suppliers that have split them, and against free. The models that survive that arithmetic either collect where costs are low and sell where demand is, or sell something whose price is not set per hour.
4 · What became scarce
Three inputs replaced volume as the constraint, and none of them is produced by collecting more hours.
Positions are the author’s qualitative assessment of structural properties, not a score of individual companies. Rights transferability means what a buyer can be granted, not how ethically the data was gathered — several models score low here while operating to a high ethical standard.
Rights that transfer. An individual can license their own likeness. They cannot license their employer’s premises, the processes visible in frame, or the colleagues who appear alongside them, and most employment agreements restrict on-site recording before consent is even reached. A rights chain that stops at the contributor cannot grant what workplace footage requires, however carefully it was gathered.
Refinement to a trainable object. Video is not a training set. Per-joint contact state, metric-scaled geometry and retargeted trajectories validated against a specific end effector are the difference between a corpus and a policy input, and they are the layers open releases do not carry.
Evidence of lift. The category has almost no independent evaluation. Buyers are asked to believe that a corpus improves a policy, and the supplier that sells the data is usually also the one that measures it. Neutral evaluation is the layer most obviously missing and the one likeliest to appear next.
5 · The comparison
Nine suppliers that can deliver egocentric hours today. One rule: every row carries a source, and no row is included without one. MOVAS AI appears as one supplier among several and is not scored above the others.
| Provider | Supply model | Published scale | Annotation depth | Rights structure | Delivery | Typical use case | Source |
|---|---|---|---|---|---|---|---|
| MOVAS AI | Employer-signed third-party workplaces | 100,000 h in inventory | L1–L5 control annotation + 4D + MoCap + 3D + retargeting to 5 end effectors | Employer agreement, then individual consent; per-clip consent id | UMI · LeRobot · RDT · OXE · native | Contact-rich work where workplace rights must survive legal review | movas.ai — self-reported |
| Maxinsights | Global capture network plus in-house processing | 1.5M h held · 2M h delivered to date | Multi-layer incl. tactile and force; recovery of occlusion and fast-motion cases | Not published | Not published | Deepest published annotation stack at volume | Xinzhiyuan, 11 Sep 2026 — company-stated |
| Build AI | Factory workers on proprietary head-mounted rigs; open release | 100,405 h free (Apache 2.0) · 14,228 workers · 10.8B frames | 256p raw video — no IMU, no hand pose, no contact labels | Apache 2.0, free commercial use | Hugging Face streaming | Free industrial volume for visual pre-training | Hugging Face dataset card |
| Lightwheel | Real-world capture plus simulation | >1.5M h human data delivered · 25,000+ environment nodes · 100,000+ task types | Video, audio, depth, structured labels | Not published | Sim-ready, robot-agnostic | Breadth across task types, with a simulation path | Xinhua Finance, 25 Jun 2026 |
| Mecka | Paid contributor network, body sensors and phones | Hours not published | Human motion capture, oriented to world models | Not published | Not published | World-model programmes needing human motion at volume | Fortune, 1 Jun 2026; TechCrunch, 11 Sep 2026 |
| XDOF | Teleoperation and wearable sensing; in-house pipelines | Hours not published · 130,000+ teleop episodes open-sourced | Not published | Not published | Not published | Robot-native trajectories rather than human video | TechCrunch, 17 Jun 2026 |
| Human Archive | Indian home-services, hotel and restaurant workers in camera caps | 1,000+ headsets deployed · base rate $1/hour | Not published | Individual consent plus hourly compensation | Not published | Low-cost volume in service settings | TechCrunch, 26 May 2026 |
| Appen | Enterprise contributor network | 1M+ contributors | Boxes, segmentation, action labels | Enterprise governance framework | Enterprise pipelines | Compliance-heavy enterprise programmes | appen.com — self-reported |
“Not published” means the provider does not state it publicly, not that it is absent. “Self-reported” means published by the provider and not independently verified.
Held out of the table
Awign
No primary source located. The figure traces to industry round-ups rather than the company or named coverage.
Luel
The only traceable source is an SEO round-up published by an annotation vendor that ranks itself first.
EgoData
Publishes no scale figure. Its sample portal is client-rendered and returns no indexable content.
6 · What to watch
Three developments would change the analysis above, in descending order of likelihood.
Questions this raises
No. It has fragmented into five supply models that are not substitutes for one another: open corpora, gig networks, contributor networks with deep annotation, platform and OEM in-house collection, and employer-signed third-party capture. A buyer choosing between them is choosing between different legal objects, not different prices for the same thing.
Because a permissively licensed release reached the size of a large commercial corpus. Build AI’s Egocentric-100K carries 100,405 hours under Apache 2.0 with free commercial use. Once free supply matches paid inventory, the paid hour has to be different rather than merely more numerous.
They are in different places. Booked revenue and the named frontier buyers concentrate in North America; collection capacity and new capital concentrate in Asia, with China holding the largest operated capacity and India the lowest unit cost. Any supplier whose cost base and revenue base sit in the same region is competing against one that has split them.
Three things. The right to use footage recorded inside someone else’s operating business; the refinement that converts video into a trajectory a specific robot can execute; and independent evidence that the data moved a policy’s success rate. None of the three is produced by collecting more hours.
Probably, and that is the wrong thing to watch. The more interesting development is that open releases have started carrying manipulation-relevant annotation rather than raw video alone, with commercial suppliers contributing to shared platforms. If annotation follows hours into the commons, the scarce layer moves again — towards rights and proof.