Self-Bootstrapping Frozen VLAs on Empirically Zero-Success Layouts
Abstract
Frozen vision-language-action (VLA) policies can be highly reliable near the object layouts seen during training, yet fail under larger spatial rearrangements. Across six LIBERO-Goal tasks and 90 unseen layouts, the libero-goal OpenVLA-OFT checkpoint succeeds on only 5% of trials over five evaluation seeds, with all successes at displacements below 15 cm. We call this failure pattern an empirical zero-success regime. Methods that learn from the policy's own successes—self-improvement and RL fine-tuning—have no positive signal on these layouts, while human demonstrations must be re-collected whenever spatial coverage is extended. We instead take a single successful policy rollout per task, geometrically transport it to new object configurations using only the object poses at the initial step (t = 0), and re-execute the transported waypoints under closed-loop tracking, retaining executions that satisfy both the task predicate and a final-position criterion—no human input, no reward, and no synthesized observations. From 806 attempts, this procedure yields 600 verified demonstrations. A single 2000-step LoRA fine-tuning stage on these demonstrations raises success from 5% to 82% over all 90 layouts, and from 0% to 79% on empirically zero-success layouts. In our ablation, policy references yield more accepted demonstrations than LIBERO human demonstrations under the same generation conditions; with training data matched in layouts and count, policy references remain slightly ahead after fine-tuning (+3.1 pp) and better preserve success on the original LIBERO layouts (81.7% versus 71.7%, consistent across both training seeds). In our setting, successful rollouts produced by the policy itself thus provide an effective reference for extending spatial coverage without human demonstrations.