First-person video, 4D multi-view capture, expert annotation and open evaluation — the complete data stack for training embodied AI, VLA models and world models.
The largest commercially-licensed egocentric corpus assembled for Physical AI. Captured head-mounted, in live working environments, with electronic image stabilisation disabled so ego-motion stays recoverable downstream.
Nine scene families weighted toward contact-rich commercial, domestic and industrial work — the environments that produce the highest density of bimanual, long-horizon manipulation per hour.
A multimodal data foundation built for spatial reasoning and physical-world understanding — at the scale required to train next-generation Physical AI.
100,000 hours of head-mounted, first-person video across commercial, domestic, industrial, food-service and retail environments — engineered for VLA pre-training and world-model learning.
Explore dataset10,000+ hours of time-synchronised multi-camera capture that resolves single-view occlusion and recovers true 3D spatial structure across time.
Explore datasetReconstructed object meshes and full environment scans, physically scaled and simulation-ready for Isaac Sim, MuJoCo and Genesis.
Explore datasetOptical and inertial motion capture of full-body and hand kinematics during real manipulation tasks, aligned to the egocentric and multi-view corpora.
Explore datasetVideo, audio, text and action streams delivered as hardware-synchronised bundles that drop directly into a training loop without alignment work.
Explore datasetAn annotation layer spanning object detail, camera dynamics, scene summary, visual properties and semantic context — graded by domain experts, not crowd workers.
Explore datasetAn industrial-grade end-to-end platform for ingestion, high-density annotation, curation and alignment evaluation — designed for complex multimodal assets.
LLM-augmented preprocessing with auto-segmentation and multi-camera tracking. Cuts manual annotation time by up to 70% versus standard pipelines.
Nested schemas mapping visual features directly to physical properties, with dynamic hierarchies and cross-task taxonomies.
Automated PII scrubbing, cryptographic watermarking and full IP audit trails. Air-gapped deployment available; ISO 27001 aligned.
Purpose-built data architectures for the hardest problems in Physical AI.
From perceiving the world to stably executing complex tasks.
Producing the real-world data that cannot otherwise be collected.
Making AI adaptable to complex, dynamic industrial environments.
Movas Embodied Video Benchmark — a shared yardstick for the embodied AI community. Core tasks, evaluation code and the public leaderboard are openly available.
12 core manipulation tasks with an evaluation harness and public leaderboard, released under a permissive licence.
Standardised protocols, sealed test splits and multi-embodiment support for fair cross-team comparison.
Optional commercial tier: extended task suites, private holdout sets and SLA-backed evaluation infrastructure.
Built at the intersection of data infrastructure and Physical AI — designed to address the structural limits of generic crowdsourced annotation.
Annotations validated by roboticists, kinematics researchers and spatial-AI specialists to ensure physical accuracy.
Proprietary sync-and-reconstruction pipeline for massive-scale egocentric and 4D multi-camera collection.
Up to 70% reduction in manual annotation time via automated segmentation, multi-camera tracking and LLM-assisted pipelines.
Egocentric video is captured from a head-mounted camera worn by a person performing a task, so the viewpoint approximates what a humanoid robot would see doing the same work. Because human and humanoid kinematic chains are similar, policies pre-trained on egocentric video transfer to robots with far less adaptation than third-person video requires — and it is roughly a third to a fifth the cost of teleoperated real-robot collection.
100,000 hours across nine scene families and eighteen sub-classes. Commercial operations, home and domestic work, and food preparation are the three largest, together making up roughly 63 percent. The full breakdown is published on the MOVAS-Ego 100K dataset page.
UMI, LeRobot, RDT and Open-X-Embodiment, plus a native MOVAS format that preserves the full annotation stack. Data can be delivered to your cloud, to an air-gapped environment, or accessed through the Movas-OS platform.
Yes. Contributors sign employer-mediated agreements that cover commercial training use, and the corpus carries full provenance and audit records. This is the practical difference between MOVAS and academic egocentric corpora, which usually carry research-only terms.
Ready to fuel your model? Partner with MOVAS AI for proprietary Physical AI datasets and infrastructure.
Book a Demo