RecoverBench: A Systematic Benchmark for Error Recovery in Robotic Manipulation
Abstract
Robotic manipulation policies have achieved remarkable success in controlled settings, but they catastrophically fail when execution errors occur---dropped objects, collisions, misaligned grasps. Despite the critical importance of error recovery for real-world deployment, no existing benchmark systematically evaluates manipulation policies' ability to recover from such errors. We introduce RecoverBench, the first systematic benchmark for error recovery in robotic manipulation. We propose a taxonomy of 12 Error Skills organized into 5 Recovery Behavior Groups (RBGs), spanning 24 error subtypes across 2 difficulty degrees. Our Error Skill pipeline automatically generates 1,360 reproducible error scenes from clean demonstration trajectories across 6 manipulation tasks. Each error scene includes complete simulation state snapshots, environment fingerprints, and RNG states for deterministic reproduction. Beyond evaluation, RecoverBench also supports recovery-demo collection and augmentation, and we show that this recovery supervision improves recovery performance. We evaluate 4 representative manipulation policies on RecoverBench. Even the best-performing policy reaches only 48.6% average recovery success, ranging from 32.0% on threading to 76.7% on stack, showing that current methods still lack robust recovery. These results highlight the need for dedicated recovery evaluation and position RecoverBench as a benchmark for improving recovery robustness in manipulation policies. Code, data, and evaluation protocols will be released.