MOVAS-Ego 100K — Egocentric Video Dataset for Physical AI
100,000 hours of head-mounted first-person video captured during professional work — every hour under an employer-signed agreement, with hand pose, contact and retargeted trajectories aligned on the same clip.
Overview
MOVAS-Ego 100K is a leading commercially-licensed egocentric corpus assembled specifically for Physical AI. Every hour is captured from a head-mounted rig worn by a worker performing real, unscripted tasks in a live environment — not staged demonstrations and not scraped web video.
Electronic image stabilization and lens-distortion correction are disabled at capture time. This is deliberate: EIS silently rewrites the relationship between camera motion and scene motion, destroying the ego-motion signal that downstream SLAM, hand-pose and retargeting stages depend on. It cannot be corrected after the fact.
The corpus is weighted toward contact-rich, bimanual, long-horizon work. Food preparation, contact-rich assembly and line work, fine assembly and specialized services together account for 31,428 hours of contact-rich work — the slice where manipulation policies fail first, and the slice that generic crowdsourced or web-scraped video under-represents.
Scene Distribution
The corpus is deliberately unbalanced. Commercial and domestic environments lead on hours because that is where MOVAS holds the most employer agreements and the longest site tenure. They are not the contact-dense end of the corpus — that slice sits in food preparation, industrial work and fine assembly. Industrial operations were added as a dedicated family to serve line-side and warehouse policies, which have materially different object statistics from retail.
- Commercial Operations 25.6%
- Food Preparation 17.8%
- Domestic & Care Services 17.2%
- Industrial & Production 11%
- General Retail 9.8%
- F&B Storefront 7.5%
- Specialized Services 5.4%
- Handcraft & Fine Assembly 3.5%
- Market, Agri & Mobility 2.2%
Full 18 sub-class breakdown
| Sub-Class | Scene Family | Duration | Share |
|---|---|---|---|
| Storefront Service & Counter Operations | Commercial Operations | 25,636 h | 25.64% |
| Food Preparation & Cooking | Food Preparation | 17,770 h | 17.77% |
| Shelf Replenishment & Inventory | General Retail | 9,823 h | 9.82% |
| Housekeeping & Tidying | Domestic & Care Services | 8,200 h | 8.2% |
| Order, Plating & Table Service | F&B Storefront | 7,466 h | 7.47% |
| Kitchen & Dishwashing | Domestic & Care Services | 5,100 h | 5.1% |
| Assembly & Line Work | Industrial & Production | 4,800 h | 4.8% |
| Laundry & Textile Handling | Domestic & Care Services | 3,900 h | 3.9% |
| Sorting & Packing | Industrial & Production | 3,600 h | 3.6% |
| Handcraft & Fine Assembly | Handcraft & Fine Assembly | 3,461 h | 3.46% |
| Home Repair & Appliance Service | Specialized Services | 2,610 h | 2.61% |
| Warehouse & Material Handling | Industrial & Production | 2,600 h | 2.6% |
| Wet Market & Bazaar | Market, Agri & Mobility | 1,580 h | 1.58% |
| Professional Repair | Specialized Services | 1,523 h | 1.52% |
| Automotive Service | Specialized Services | 1,036 h | 1.04% |
| Agriculture & Farming | Market, Agri & Mobility | 356 h | 0.36% |
| Transportation & Mobility | Market, Agri & Mobility | 311 h | 0.31% |
| Beauty & Personal Care | Specialized Services | 228 h | 0.23% |
| Total | 100,000 h | 100% |
Proportions derived from the MOVAS internal data catalog and normalized to the 100,000-hour production corpus. Sub-classes are grouped into nine top-level scene families; licensing is available at either level.
Data Structure
Directory layout as delivered. Every clip is self-contained: no cross-referencing required to train on a single sample.
movas-ego-100k/
├── clips/
│ └── {scene_family}/{clip_uuid}/
│ ├── video_left.mp4 / video_right.mp4 # 1080p · 30 fps nominal, H.265, bit-exact remux, EIS off
│ ├── imu.npz # 476 Hz 6-axis, hw-synced
│ └── meta.json # scene, duration, rig, consent id
├── hand/
│ └── {clip_uuid}/
│ ├── mano_left.npz # [T, 51] MANO params
│ ├── mano_right.npz
│ └── contact.json # per-frame, per-joint contact
├── objects/
│ └── {clip_uuid}/
│ ├── {obj_id}_pose.npz # [T, 4, 4] 6-DoF
│ └── {obj_id}_mask.npz
├── retarget/
│ └── {clip_uuid}/
│ ├── shadow_hand.npz
│ ├── inspire_hand.npz
│ ├── allegro.npz
│ └── franka_gripper.npz
├── annotations/
│ └── {clip_uuid}.json # Control L1-L5 label stack
└── manifest.parquet # index + per-clip quality score
Annotation Schema
Task
Task identity and success/failure outcome.
Phase
Segmented action phases: approach, pre-grasp, grasp, transport, manipulate, release.
Object
Open-vocabulary object identity, material and state transitions.
Keyframe
1 fps dense keyframes plus every phase-boundary frame.
Contact
Per-hand, per-joint contact points and duration derived from MANO surface distance.
Quality Gates
Published thresholds, enforced at ingest. Records that fail are rejected or flagged in the manifest — never silently included.
| Gate | Threshold | Method |
|---|---|---|
| Hand visibility | ≥ 50% of frames | Clips below threshold are rejected at ingest. |
| IMU–video drift | < 1 ms / 60 s | Hardware timestamp sync, verified per clip. |
| Reprojection error | < 3 px @ 1080p | Geometric consistency gate after 4D reconstruction. |
| Composite score | ≥ 70 / 100 | Weighted across reconstruction, hand pose, tracking, retarget feasibility, sim replay. |
Formats & Access
Export formats
Modalities included
Cloud delivery
Direct to your S3, GCS or OSS bucket. Manifest-driven incremental sync.
Air-gapped
Physical media transfer for regulated or offline environments.
Via Movas-OS
Stream and re-annotate in place through the Movas-OS platform.
Use Cases
VLA pre-training
Large-scale visuomotor pre-training where scale and action diversity dominate.
World-model pre-training
Next-frame and next-state prediction grounded in real physical dynamics, with recoverable ego-motion.
Human-to-robot transfer
Retargeted trajectories give direct supervision for dexterous and parallel-jaw end effectors.
Affordance & contact learning
L5 contact annotations supervise where and how objects are grasped.
Related Datasets
What the sensor actually wrote.
Measured on delivered assets, not quoted from a device datasheet.
Video
Optics and calibration
Inertial
Integrity
Figures are the measured range across delivered assets. Per-asset values ship on the data card and on every sample record. Why each threshold sits where it does: MV-RSCH-010. What the rejected half contains: MV-RSCH-008. How end-to-end yield is computed: MV-RSCH-002.