Benchmarking Robotic Manipulation under Physically Realistic Dynamics
Abstract
Object manipulation in the dynamic real world requires understanding of how the environment evolves and how to respond to it accordingly. The existing benchmarks, however, do not adequately address these needs. The physics-based benchmarks focus primarily on predictive capability of models, while omitting action component. The dynamic robotic manipulation benchmarks contain only basic physically implausible object motions and evaluate policies synchronously, so physical anticipation is never necessary and inference cost is free. There remains a gap for evaluating robotics policies' both predictive reasoning and actions in physically realistic dynamic environments. To this end, we propose RoboDyna, a benchmark for evaluating physical understanding and actions of manipulation policies in dynamic scenes. RoboDyna consists of 20 conceptual dynamic tasks with simplified objects and minimum scene clutter in five categories under a two-factor difficulty design for evaluating core physical understanding and capability of policies. In addition, 10 complex household tasks with realistic scenes are designed to measure how these skills can be transferred to real-world scenarios. Our benchmark contains 3500 expert demonstrations and features protocols for synchronized (frozen during inference time) and asynchronous evaluation, success and partial metrics to measure progress, and interactive tools for collecting human data and establishing human baseline performance under teleoperation and scene-interactive regimes. Our extensive evaluation of state-of-the-art VLA and WAM policies on RoboDyna reveals that prediction grants the WAM the ability to anticipate the consequences of its own actions in initially static scenes, but under multi-object motion and complex patterns it fails in both prediction and action. Comparison between synchronous and asynchronous evaluation shows that inference latency changes which rollouts succeed more than how many, with systematic losses on time-critical tasks. Further evaluation shows policies generalize to unseen compositions of seen dynamics. The code, benchmark, and complete results will be released upon acceptance.