LIBERO-MAX: Do Robot Policies Adapt When the World Changes?
Abstract
Simulation benchmarks usually fix exogenous conditions at reset, although a deployed robot may encounter changes after it begins acting. We introduce LIBERO-MAX, an evaluation with 8,000 Base/Dynamic pairs that share the task, initial state, policy seed, and executed action prefix; only the Dynamic rollout then receives one persistent event. Restricting each pair to one event isolates the change's incremental effect on task completion, while recurrent and compound events remain a separate evaluation target. Across fourteen VLA, hybrid, and world-action policies, success falls by 11.0 to 25.7 points, and 20.8% to 56.1% of episodes solved in Base fail after the event. Target and receptacle relocation and camera and sensor changes produce the largest shared losses; changing query cadence over a fixed 800-pair sweep does not remove the gap. We also release LIBERO-MAX Lite, a fixed 800-pair subset that estimates every reported rate within 2.4 points and preserves 88 of 91 pairwise Dynamic orderings, enabling faster development of policies that respond to changes during execution.