First-Person Video for Embodied AI

MOVAS-Ego 100K — Egocentric Video Dataset for Physical AI

100,000 hours of head-mounted, first-person video across commercial, domestic, industrial, food-service and retail environments — engineered for VLA pre-training and world-model learning.

Total Duration 100,000+ hrs
Scene Families 9 / 18 sub-classes
Capture Rig Head-mounted, EIS-off
Frame Rate 30 fps @ 1080p

Overview

MOVAS-Ego 100K is the largest commercially-licensed egocentric corpus assembled specifically for Physical AI. Every hour is captured from a head-mounted rig worn by a worker performing real, unscripted tasks in a live environment — not staged demonstrations and not scraped web video.

Electronic image stabilisation and lens-distortion correction are disabled at capture time. This is deliberate: EIS silently rewrites the relationship between camera motion and scene motion, destroying the ego-motion signal that downstream SLAM, hand-pose and retargeting stages depend on. Correcting it after the fact is not possible.

The corpus is weighted toward commercial and food-service work because those environments produce the highest density of contact-rich, bimanual, long-horizon manipulation per hour — the exact distribution that generic crowdsourced or web-scraped video under-represents.

100,000+ hrs
Total Duration
9 / 18 sub-classes
Scene Families
Head-mounted, EIS-off
Capture Rig
30 fps @ 1080p
Frame Rate

Scene Distribution

The corpus is deliberately unbalanced. Commercial and domestic environments lead because they generate the highest density of contact-rich, bimanual, long-horizon manipulation per recorded hour, and because they are the two settings where humanoid deployment is nearest term. Industrial operations were added as a dedicated family to serve line-side and warehouse policies, which have materially different object statistics from retail.

Commercial Operations — 25.6% Home & Domestic — 19.8% Food Preparation — 17.8% Industrial Operations — 11% General Retail — 9.8% F&B Storefront — 7.5% Handcraft & DIY — 3.5% Specialised Services — 2.8% Market, Agri & Mobility — 2.2% 100K HOURS
  • Commercial Operations 25.6%
  • Home & Domestic 19.8%
  • Food Preparation 17.8%
  • Industrial Operations 11%
  • General Retail 9.8%
  • F&B Storefront 7.5%
  • Handcraft & DIY 3.5%
  • Specialised Services 2.8%
  • Market, Agri & Mobility 2.2%

Full 18 sub-class breakdown

MOVAS-Ego 100K scene class distribution by duration
Sub-ClassScene FamilyDurationShare
Storefront Service & Counter OperationsCommercial Operations25,636 h25.64%
Food Preparation & CookingFood Preparation17,770 h17.77%
Shelf Replenishment & InventoryGeneral Retail9,823 h9.82%
Home Living & TidyingHome & Domestic8,200 h8.2%
Order, Plating & Table ServiceF&B Storefront7,466 h7.47%
Kitchen & DishwashingHome & Domestic5,100 h5.1%
Assembly & Line WorkIndustrial Operations4,800 h4.8%
Laundry & Fabric HandlingHome & Domestic3,900 h3.9%
Sorting & PackingIndustrial Operations3,600 h3.6%
Handcraft & Fine AssemblyHandcraft & DIY3,461 h3.46%
Home DIY & MaintenanceHome & Domestic2,610 h2.61%
Warehouse & Material HandlingIndustrial Operations2,600 h2.6%
Wet Market & BazaarMarket, Agri & Mobility1,580 h1.58%
Professional RepairSpecialised Services1,523 h1.52%
Automotive ServiceSpecialised Services1,036 h1.04%
Agriculture & FarmingMarket, Agri & Mobility356 h0.36%
Transportation & MobilityMarket, Agri & Mobility311 h0.31%
Beauty & Personal CareSpecialised Services228 h0.23%
Total100,000 h100%

Proportions derived from the MOVAS internal data catalogue and normalised to the 100,000-hour production corpus. Sub-classes are grouped into nine top-level scene families; licensing is available at either level.

Data Structure

Directory layout as delivered. Every clip is self-contained: no cross-referencing required to train on a single sample.

movas-ego-100k/
├── clips/
│   └── {scene_family}/{clip_uuid}/
│       ├── video.mp4              # 1080p30, H.264 CRF 18, EIS off
│       ├── imu.npz                # 200 Hz 9-axis, hw-synced
│       └── meta.json              # scene, duration, rig, consent id
├── hand/
│   └── {clip_uuid}/
│       ├── mano_left.npz          # [T, 51] MANO params
│       ├── mano_right.npz
│       └── contact.json           # per-frame binary contact
├── objects/
│   └── {clip_uuid}/
│       ├── {obj_id}_pose.npz      # [T, 4, 4] 6-DoF
│       └── {obj_id}_mask.npz
├── retarget/
│   └── {clip_uuid}/
│       ├── shadow_hand.npz
│       ├── inspire_hand.npz
│       ├── allegro.npz
│       └── franka_gripper.npz
├── annotations/
│   └── {clip_uuid}.json           # L1-L5 label stack
└── manifest.parquet               # index + per-clip quality score

Annotation Schema

L1

Task

High-level task identity and success/failure outcome.

L2

Phase

Segmented action phases: approach, pre-grasp, grasp, transport, manipulate, release.

L3

Object

Open-vocabulary object identity, material, and state transitions.

L4

Keyframe

1 fps dense keyframes plus every phase-boundary frame.

L5

Contact

Per-hand, per-joint contact points and contact duration derived from MANO surface distance.

Quality Gates

Published thresholds, enforced at ingest. Records that fail are rejected or flagged in the manifest — never silently included.

GateThresholdMethod
Hand visibility≥ 50% of framesClips below threshold are rejected at ingest.
IMU–video drift< 1 ms / 60 sHardware timestamp sync, verified per clip.
Reprojection error< 3 px @ 1080pGeometric consistency gate after 4D reconstruction.
Composite score≥ 70 / 100Weighted across reconstruction, hand pose, tracking, retarget feasibility, sim replay.

Formats & Access

Export formats

UMILeRobotRDTOpen-X-EmbodimentMOVAS-Ego native

Modalities included

RGB videoIMU 200 HzHand pose (MANO)Contact statesTask & phase labelsRetargeted trajectories

Cloud delivery

Direct to your S3, GCS or OSS bucket. Manifest-driven incremental sync.

Air-gapped

Physical media transfer for regulated or offline environments.

Via Movas-OS

Stream and re-annotate in place through the Movas-OS platform.

Use Cases

VLA pre-training

Large-scale visuomotor pre-training where scale and action diversity dominate.

World-model learning

Next-frame and next-state prediction grounded in real physical dynamics.

Human-to-robot transfer

Retargeted trajectories give direct supervision for dexterous and parallel-jaw end effectors.

Affordance & contact learning

L5 contact annotations supervise where and how objects are grasped.

Related Datasets

FAQ

MOVAS-Ego 100K — Frequently Asked Questions

Ego4D and Ego-Exo4D are research corpora built for perception benchmarks — activity recognition, episodic memory, and captioning. MOVAS-Ego 100K is built for control. Every clip carries hand pose, contact state, and trajectories already retargeted onto common end effectors, and capture-side EIS is disabled so ego-motion remains recoverable. It is also cleared for commercial training use.

Yes, at the firmware level on every rig. EIS applies a per-frame homography that decouples pixel motion from true camera motion. Any downstream stage that estimates ego-motion — SLAM, hand-pose lifting, object 6-DoF tracking — inherits that error, and it cannot be inverted after the fact.

Five-finger anthropomorphic hands (Shadow, Inspire), four-finger hands (Allegro), parallel-jaw grippers (Franka, UR5), and three-finger underactuated hands (Robotiq 3F). Additional embodiments can be generated on request from the source MANO parameters.

Every contributor signs an employer-mediated agreement covering commercial training use. Faces, licence plates, screens and identity documents are blurred on-device before upload; precise GPS is discarded at the edge and only city-level geography is retained. Pre-anonymisation source material is destroyed on a 90-day schedule.

Yes. The corpus is licensed by scene family and by annotation depth. A typical first engagement is a single scene family at L1–L3 annotation, expanding to L5 and additional families once training value is demonstrated.

Request Dataset Access

Send us the task you are training for and we will scope the smallest subset that moves your metric.

Book a Demo