MOVAS-Ego 100K — Egocentric Video Dataset for Physical AI
100,000 hours of head-mounted, first-person video across commercial, domestic, industrial, food-service and retail environments — engineered for VLA pre-training and world-model learning.
Overview
MOVAS-Ego 100K is the largest commercially-licensed egocentric corpus assembled specifically for Physical AI. Every hour is captured from a head-mounted rig worn by a worker performing real, unscripted tasks in a live environment — not staged demonstrations and not scraped web video.
Electronic image stabilisation and lens-distortion correction are disabled at capture time. This is deliberate: EIS silently rewrites the relationship between camera motion and scene motion, destroying the ego-motion signal that downstream SLAM, hand-pose and retargeting stages depend on. Correcting it after the fact is not possible.
The corpus is weighted toward commercial and food-service work because those environments produce the highest density of contact-rich, bimanual, long-horizon manipulation per hour — the exact distribution that generic crowdsourced or web-scraped video under-represents.
Scene Distribution
The corpus is deliberately unbalanced. Commercial and domestic environments lead because they generate the highest density of contact-rich, bimanual, long-horizon manipulation per recorded hour, and because they are the two settings where humanoid deployment is nearest term. Industrial operations were added as a dedicated family to serve line-side and warehouse policies, which have materially different object statistics from retail.
- Commercial Operations 25.6%
- Home & Domestic 19.8%
- Food Preparation 17.8%
- Industrial Operations 11%
- General Retail 9.8%
- F&B Storefront 7.5%
- Handcraft & DIY 3.5%
- Specialised Services 2.8%
- Market, Agri & Mobility 2.2%
Full 18 sub-class breakdown
| Sub-Class | Scene Family | Duration | Share |
|---|---|---|---|
| Storefront Service & Counter Operations | Commercial Operations | 25,636 h | 25.64% |
| Food Preparation & Cooking | Food Preparation | 17,770 h | 17.77% |
| Shelf Replenishment & Inventory | General Retail | 9,823 h | 9.82% |
| Home Living & Tidying | Home & Domestic | 8,200 h | 8.2% |
| Order, Plating & Table Service | F&B Storefront | 7,466 h | 7.47% |
| Kitchen & Dishwashing | Home & Domestic | 5,100 h | 5.1% |
| Assembly & Line Work | Industrial Operations | 4,800 h | 4.8% |
| Laundry & Fabric Handling | Home & Domestic | 3,900 h | 3.9% |
| Sorting & Packing | Industrial Operations | 3,600 h | 3.6% |
| Handcraft & Fine Assembly | Handcraft & DIY | 3,461 h | 3.46% |
| Home DIY & Maintenance | Home & Domestic | 2,610 h | 2.61% |
| Warehouse & Material Handling | Industrial Operations | 2,600 h | 2.6% |
| Wet Market & Bazaar | Market, Agri & Mobility | 1,580 h | 1.58% |
| Professional Repair | Specialised Services | 1,523 h | 1.52% |
| Automotive Service | Specialised Services | 1,036 h | 1.04% |
| Agriculture & Farming | Market, Agri & Mobility | 356 h | 0.36% |
| Transportation & Mobility | Market, Agri & Mobility | 311 h | 0.31% |
| Beauty & Personal Care | Specialised Services | 228 h | 0.23% |
| Total | 100,000 h | 100% |
Proportions derived from the MOVAS internal data catalogue and normalised to the 100,000-hour production corpus. Sub-classes are grouped into nine top-level scene families; licensing is available at either level.
Data Structure
Directory layout as delivered. Every clip is self-contained: no cross-referencing required to train on a single sample.
movas-ego-100k/
├── clips/
│ └── {scene_family}/{clip_uuid}/
│ ├── video.mp4 # 1080p30, H.264 CRF 18, EIS off
│ ├── imu.npz # 200 Hz 9-axis, hw-synced
│ └── meta.json # scene, duration, rig, consent id
├── hand/
│ └── {clip_uuid}/
│ ├── mano_left.npz # [T, 51] MANO params
│ ├── mano_right.npz
│ └── contact.json # per-frame binary contact
├── objects/
│ └── {clip_uuid}/
│ ├── {obj_id}_pose.npz # [T, 4, 4] 6-DoF
│ └── {obj_id}_mask.npz
├── retarget/
│ └── {clip_uuid}/
│ ├── shadow_hand.npz
│ ├── inspire_hand.npz
│ ├── allegro.npz
│ └── franka_gripper.npz
├── annotations/
│ └── {clip_uuid}.json # L1-L5 label stack
└── manifest.parquet # index + per-clip quality score
Annotation Schema
Task
High-level task identity and success/failure outcome.
Phase
Segmented action phases: approach, pre-grasp, grasp, transport, manipulate, release.
Object
Open-vocabulary object identity, material, and state transitions.
Keyframe
1 fps dense keyframes plus every phase-boundary frame.
Contact
Per-hand, per-joint contact points and contact duration derived from MANO surface distance.
Quality Gates
Published thresholds, enforced at ingest. Records that fail are rejected or flagged in the manifest — never silently included.
| Gate | Threshold | Method |
|---|---|---|
| Hand visibility | ≥ 50% of frames | Clips below threshold are rejected at ingest. |
| IMU–video drift | < 1 ms / 60 s | Hardware timestamp sync, verified per clip. |
| Reprojection error | < 3 px @ 1080p | Geometric consistency gate after 4D reconstruction. |
| Composite score | ≥ 70 / 100 | Weighted across reconstruction, hand pose, tracking, retarget feasibility, sim replay. |
Formats & Access
Export formats
Modalities included
Cloud delivery
Direct to your S3, GCS or OSS bucket. Manifest-driven incremental sync.
Air-gapped
Physical media transfer for regulated or offline environments.
Via Movas-OS
Stream and re-annotate in place through the Movas-OS platform.
Use Cases
VLA pre-training
Large-scale visuomotor pre-training where scale and action diversity dominate.
World-model learning
Next-frame and next-state prediction grounded in real physical dynamics.
Human-to-robot transfer
Retargeted trajectories give direct supervision for dexterous and parallel-jaw end effectors.
Affordance & contact learning
L5 contact annotations supervise where and how objects are grasped.