Benchmarking Human-Video-Prompted Robot Adaptation under Deployment Scene Constraints
Abstract
Real-world constraints in embodied deployment are not only computational: the physical environment itself restricts how a task can be executed. A household robot cannot assume that the trajectory, object instance, layout, or even functional tool shown in demonstrations remains available in its own scene. We study how human-video-prompted models perform under such scene-imposed constraints. The video specifies the task once; the robot must infer from observations about how to realize the task demonstrated under its own scene. We introduce The Imitator Game, a four-level benchmark that progressively strengthens demonstration-execution constraints from trajectory-preserving scenes to functional substitution; IG-10K, an environment-aligned paired human-robot dataset with 20,000+ paired episodes, spanning 200+ task variants, collected in both simulation and real world; and Imitator Arena, which combines automated simulation metrics with human evaluation of anonymized rollouts. Across nine state-of-the-art methods, models can indeed learn real-time execution from human demonstration, with video-prompted models performing better. At the same time, pretraining makes most policies easier to adapt to novel tasks using only 10 few-shot demonstrations, with gains that generally increase with pretraining scale. On real hardware, models may adapt object layout and appearance constraints but fail when the deployment requires intent-level transfer with affordance constraints. The results show that on-device, human-prompted robot deployment is achievable at scale and the prompting methods matter, while the deployment scene constraint largely affects the model performance. The project website and access to Imitator Arena are available at https://imitator-game.github.io.