Restore, Don’t Retry: Diagnosing and Recovering from Observation-Space Shift at Long-Horizon Skill Seams
Pranav Wagh ⋅ Yu Fang ⋅ Yue Yang ⋅ Mingyu Ding
Abstract
Long-horizon manipulation is often built by chaining individually trained skills, yet such chains collapse at the seams where one skill hands off to the next. The cause is Observation-Space Shift (OSS), the failure identified by the BOSS benchmark: skill {i} is executed from the terminal state of skill {i}-1, which lies off skill {i}'s behavior-cloning training support. On BOSS-44, full-chain success across two seams drops from $0.885$ to $0.076$ on the ten most affected chains. By restoring one scene component at a time, we localize the shift to displaced drawer articulation and secondary objects rather than to the robot's pose or the manipulated object, and show that the two components interact: restoring both together recovers the full privileged-oracle gain, whereas restoring either alone does not. We then show, through a suite of controls, that none of the standard alternatives we test recovers the seam, because the shift is one of support that more sampling cannot fix. Monitor-selected test-time compute (best-of-$K$), a stronger base policy (Diffusion Policy), a video world model used as generator, monitor, and planner, and end-to-end RL recovery all fail from the same seam states. Restoration recovers the failure. We present a fully learned, non-privileged detect, restore, resume system: a task-progress monitor (a small head over a frozen DINOv3 encoder, $0.903$ pooled AUROC) with a calibrated stagnation trigger, a learned scene-restoration skill trained from a library of real failure states, and seam-robust skill fine-tuning. On ten catastrophic chains across five seeds the full system improves full-chain success $3.5\times$ over the frozen base ($0.076$ to $0.265$, $p{=}0.047$; $2.05\times$ over a budget-matched base), reaching $51\%$ of a privileged teleport oracle, using only camera images and proprioception. Finally, a real-robot study on a Franka running a served Pi0.5 policy characterizes the sim-to-real gap: under a leakage-clean cross-fit the exterior-only monitor transfers only weakly (AUROC $0.625$), limited by unlogged gripper state, and closing the loop with a language-prompted restoration (not the distilled camera-only student) recovers a fraction of otherwise-terminal failures ($3/21$, a proof of concept). The quantitative core is in simulation; the real-robot study bounds transfer under a conservative estimate and motivates wrist and gripper sensing.
Chat is not available.
Successful Page Load