Employer-Cleared Egocentric Data for Physical AI

MOVAS-Ego 100K — Egocentric Video Dataset for Physical AI

100,000 hours of head-mounted first-person video captured during professional work — every hour under an employer-signed agreement, with hand pose, contact and retargeted trajectories aligned on the same clip.

Hours in inventory 100,000+
Rights 100% employer-signed
Capture rig Head-mounted, EIS-off
Layers 6 per clip · 5 aligned frame by frame

Overview

MOVAS-Ego 100K is a leading commercially-licensed egocentric corpus assembled specifically for Physical AI. Every hour is captured from a head-mounted rig worn by a worker performing real, unscripted tasks in a live environment — not staged demonstrations and not scraped web video.

Electronic image stabilization and lens-distortion correction are disabled at capture time. This is deliberate: EIS silently rewrites the relationship between camera motion and scene motion, destroying the ego-motion signal that downstream SLAM, hand-pose and retargeting stages depend on. It cannot be corrected after the fact.

The corpus is weighted toward contact-rich, bimanual, long-horizon work. Food preparation, contact-rich assembly and line work, fine assembly and specialized services together account for 31,428 hours of contact-rich work — the slice where manipulation policies fail first, and the slice that generic crowdsourced or web-scraped video under-represents.

100,000+
Hours in inventory
100% employer-signed
Rights
Head-mounted, EIS-off
Capture rig
6 per clip · 5 aligned frame by frame
Layers

Scene Distribution

The corpus is deliberately unbalanced. Commercial and domestic environments lead on hours because that is where MOVAS holds the most employer agreements and the longest site tenure. They are not the contact-dense end of the corpus — that slice sits in food preparation, industrial work and fine assembly. Industrial operations were added as a dedicated family to serve line-side and warehouse policies, which have materially different object statistics from retail.

Commercial Operations — 25.6% Food Preparation — 17.8% Domestic & Care Services — 17.2% Industrial & Production — 11% General Retail — 9.8% F&B Storefront — 7.5% Specialized Services — 5.4% Handcraft & Fine Assembly — 3.5% Market, Agri & Mobility — 2.2% 100K HOURS
  • Commercial Operations 25.6%
  • Food Preparation 17.8%
  • Domestic & Care Services 17.2%
  • Industrial & Production 11%
  • General Retail 9.8%
  • F&B Storefront 7.5%
  • Specialized Services 5.4%
  • Handcraft & Fine Assembly 3.5%
  • Market, Agri & Mobility 2.2%

Full 18 sub-class breakdown

MOVAS-Ego 100K scene class distribution by duration
Sub-ClassScene FamilyDurationShare
Storefront Service & Counter OperationsCommercial Operations25,636 h25.64%
Food Preparation & CookingFood Preparation17,770 h17.77%
Shelf Replenishment & InventoryGeneral Retail9,823 h9.82%
Housekeeping & TidyingDomestic & Care Services8,200 h8.2%
Order, Plating & Table ServiceF&B Storefront7,466 h7.47%
Kitchen & DishwashingDomestic & Care Services5,100 h5.1%
Assembly & Line WorkIndustrial & Production4,800 h4.8%
Laundry & Textile HandlingDomestic & Care Services3,900 h3.9%
Sorting & PackingIndustrial & Production3,600 h3.6%
Handcraft & Fine AssemblyHandcraft & Fine Assembly3,461 h3.46%
Home Repair & Appliance ServiceSpecialized Services2,610 h2.61%
Warehouse & Material HandlingIndustrial & Production2,600 h2.6%
Wet Market & BazaarMarket, Agri & Mobility1,580 h1.58%
Professional RepairSpecialized Services1,523 h1.52%
Automotive ServiceSpecialized Services1,036 h1.04%
Agriculture & FarmingMarket, Agri & Mobility356 h0.36%
Transportation & MobilityMarket, Agri & Mobility311 h0.31%
Beauty & Personal CareSpecialized Services228 h0.23%
Total100,000 h100%

Proportions derived from the MOVAS internal data catalog and normalized to the 100,000-hour production corpus. Sub-classes are grouped into nine top-level scene families; licensing is available at either level.

Data Structure

Directory layout as delivered. Every clip is self-contained: no cross-referencing required to train on a single sample.

movas-ego-100k/
├── clips/
│   └── {scene_family}/{clip_uuid}/
│       ├── video_left.mp4 / video_right.mp4              # 1080p · 30 fps nominal, H.265, bit-exact remux, EIS off
│       ├── imu.npz                # 476 Hz 6-axis, hw-synced
│       └── meta.json              # scene, duration, rig, consent id
├── hand/
│   └── {clip_uuid}/
│       ├── mano_left.npz          # [T, 51] MANO params
│       ├── mano_right.npz
│       └── contact.json           # per-frame, per-joint contact
├── objects/
│   └── {clip_uuid}/
│       ├── {obj_id}_pose.npz      # [T, 4, 4] 6-DoF
│       └── {obj_id}_mask.npz
├── retarget/
│   └── {clip_uuid}/
│       ├── shadow_hand.npz
│       ├── inspire_hand.npz
│       ├── allegro.npz
│       └── franka_gripper.npz
├── annotations/
│   └── {clip_uuid}.json           # Control L1-L5 label stack
└── manifest.parquet               # index + per-clip quality score

Annotation Schema

L1

Task

Task identity and success/failure outcome.

L2

Phase

Segmented action phases: approach, pre-grasp, grasp, transport, manipulate, release.

L3

Object

Open-vocabulary object identity, material and state transitions.

L4

Keyframe

1 fps dense keyframes plus every phase-boundary frame.

L5

Contact

Per-hand, per-joint contact points and duration derived from MANO surface distance.

Quality Gates

Published thresholds, enforced at ingest. Records that fail are rejected or flagged in the manifest — never silently included.

GateThresholdMethod
Hand visibility≥ 50% of framesClips below threshold are rejected at ingest.
IMU–video drift< 1 ms / 60 sHardware timestamp sync, verified per clip.
Reprojection error< 3 px @ 1080pGeometric consistency gate after 4D reconstruction.
Composite score≥ 70 / 100Weighted across reconstruction, hand pose, tracking, retarget feasibility, sim replay.

Formats & Access

Export formats

UMILeRobotRDTOpen-X-EmbodimentMOVAS-Ego native

Modalities included

RGB video 1080p · 30 fps nominalIMU 476 Hz hardware-syncedHand pose (MANO)Per-frame contact state (inferred)Object 6-DoFRetargeted trajectories

Cloud delivery

Direct to your S3, GCS or OSS bucket. Manifest-driven incremental sync.

Air-gapped

Physical media transfer for regulated or offline environments.

Via Movas-OS

Stream and re-annotate in place through the Movas-OS platform.

Use Cases

VLA pre-training

Large-scale visuomotor pre-training where scale and action diversity dominate.

World-model pre-training

Next-frame and next-state prediction grounded in real physical dynamics, with recoverable ego-motion.

Human-to-robot transfer

Retargeted trajectories give direct supervision for dexterous and parallel-jaw end effectors.

Affordance & contact learning

L5 contact annotations supervise where and how objects are grasped.

Related Datasets

MV-SPEC-004Capture specification

What the sensor actually wrote.

Measured on delivered assets, not quoted from a device datasheet.

Video

Resolution1920 × 1080 per eye, 30 fps nominal (29.90 Hz measured)
CodecH.265 Annex-B, bit-exact remux — no re-encode, no undistortion
Stream layoutSeparate left and right files; never side-by-side
GOP≈ 30, no B-frames — frame-accurate seeking

Optics and calibration

Stereo baseline62 mm
Field of view120–160° horizontal, 84–85° vertical
Lens modelKannala-Brandt fisheye, k1–k4
Shipped with every assetPer-camera K and D, stereo extrinsics, IMU-to-camera time offset

Inertial

Rate476 Hz measured
Axes6-axis — accelerometer and gyroscope
Camera offsetSub-millisecond, written to metadata with an applied/not-applied flag
StabilizationElectronic image stabilization disabled at firmware level

Integrity

Stereo sync7 µs median; 100% of pairs under 10 ms
Frame countDelivered count equals source count — zero dropped, zero duplicates
TimestampsStrictly monotonic
AudioPCM, plus device serial per asset

Figures are the measured range across delivered assets. Per-asset values ship on the data card and on every sample record. Why each threshold sits where it does: MV-RSCH-010. What the rejected half contains: MV-RSCH-008. How end-to-end yield is computed: MV-RSCH-002.

FAQ

MOVAS-Ego 100K — Frequently Asked Questions

Ego4D and Ego-Exo4D are research corpora built for perception benchmarks — activity recognition, episodic memory and captioning. MOVAS-Ego 100K is built for control. Every clip carries hand pose, contact state and trajectories already retargeted onto common end effectors, capture-side EIS is disabled so ego-motion remains recoverable, and the corpus is cleared for commercial training use rather than research-only.

Inventory. The hours are captured, refined and in storage today. An evaluation subset ships in five business days and a full scene family in twelve. Where we quote a figure for work not yet captured — targeted re-capture, for example — we label it as such.

The operating entity of each site, through an employer agreement, with individual consent signed separately by each person filmed. An individual cannot on their own grant rights to their employer’s premises, processes or to colleagues appearing in frame. The full three-layer structure is published on our data rights page.

Yes, at firmware level on every rig. EIS applies a per-frame homography that decouples pixel motion from true camera motion. Any downstream stage that estimates ego-motion — SLAM, hand-pose lifting, object 6-DoF tracking — inherits that error, and it cannot be inverted after the fact.

Five-finger anthropomorphic hands (Shadow, Inspire), a four-finger hand (Allegro), a parallel two-finger gripper and a suction end effector — one target from each major gripper class. For other embodiments, send a URDF and we generate a custom retarget from the source MANO parameters.

Yes. The corpus is licensed by scene family and by annotation depth. A typical first engagement is a single scene family at Control L1–L3, expanding to L5 and additional families once training value is demonstrated.

Request Dataset Access

Send us the task you are training for and we will scope the smallest subset that moves your metric.

Request a Subset

Data moves robots. MOVAS moves data.