Open Benchmark

MEVB

Movas Embodied Video Benchmark — a shared yardstick for the embodied AI community. Twelve core manipulation tasks, an open evaluation harness, sealed test splits and a public leaderboard.

Apache 2.0 core12 tasks Sim + Real tracksMulti-embodiment

Three Layers

Open Core

12 core manipulation tasks with ground-truth success metrics. Evaluation harness and public leaderboard released under a permissive licence.

Reproducible Evaluation

Standardised protocols, sealed test splits and multi-embodiment support — engineered for fair, reproducible cross-team comparison.

Enterprise Extensions

Optional commercial tier for teams needing extended task suites, private holdout sets, quarterly refreshes and SLA-backed evaluation infrastructure.

Task Suite

Difficulty rises from single-hand pick-and-place through bimanual coordination and tool use to long-horizon deformable-object tasks. The final two tasks measure generalisation rather than raw capability.

IDTaskDescriptionTier
T01Cup-Pickup-PlacePick a mug from the table and place it inside a marked zone.Tier 1
T02Drawer-Open-CloseOpen a drawer fully, then close it.Tier 1
T03Door-OperationRotate a door handle and push the door open.Tier 1
T04Tray-Bimanual-TransportLift a loaded tray with both hands and move it between surfaces.Tier 2
T05Cap-Twist-OffStabilise a bottle with one hand and unscrew the cap with the other.Tier 2
T06Cloth-Fold-BimanualFold a towel twice using both hands.Tier 2
T07Sponge-WipeWipe a marked area of the table with a sponge.Tier 3
T08Spoon-Scoop-TransferTransfer granular material between bowls using a spoon.Tier 3
T09Single-Dish-WashPick up a dish, rinse, sponge, and place on the drying rack.Tier 4
T10Shirt-FoldFlatten a shirt, fold both sleeves inward, then fold from the hem.Tier 4
T11Generalization-ObjectT01 performed with an unseen mug shape, colour and size.Gen.
T12Generalization-SceneT01 performed under changed lighting, tablecloth and background.Gen.

Evaluation Protocol

Success is milestone-based and binary per milestone. A task counts as successful only when every milestone is met; partial completion is recorded separately for diagnostic use.

TaskScore = 0.50 x SuccessRate
          + 0.30 x GeneralizationScore   (L1 object + L2 scene + L3 instruction)
          + 0.20 x TimeEfficiency        (1 - actual_time / budget_time)

OverallScore = weighted mean over 12 tasks
               Tier 1 x1.0   Tier 2 x1.5   Tier 3 x2.0   Tier 4 x3.0

Splits    Train 150 / Val 25 / Test-Public 15 / Test-Hidden 10   (per task)
Tracks    Simulation 40%  +  Real-robot 60%

Built on MOVAS Data

FAQ

MEVB — Frequently Asked Questions

The core is. Twelve tasks, the evaluation harness and the public leaderboard are released under a permissive licence with no fee and no registration gate. Extended task suites, private holdout sets and SLA-backed hosted evaluation are the commercial tier.

LIBERO and CALVIN are simulation-first and tabletop-scoped. RoboFinals evaluates whole-robot systems. MEVB targets the gap between them: real physical environments, household and commercial scenes, bimanual tasks, and explicit sim-to-real comparison — evaluated at the policy level rather than the whole-machine level.

Task definitions are embodiment-agnostic. Reference implementations are provided for Franka, UR5, Unitree H1 and Optimus-class humanoids, and any embodiment can be submitted provided the milestone definitions are met.

Submissions arrive as a Docker image plus a model card. The evaluation team reproduces results independently on the sealed test split. Results that cannot be reproduced are not published.

Join the First Submission Round

MEVB v0.1 opens for submissions in 2026. Register your team to receive the harness and training split.

Book a Demo