What Makes Experience Useful for VLA Post-Training? On Informative Supervision and Reliable Evaluation
Xuan Zhao ⋅ Mohammed Elmahgiubi ⋅ Joshua J Choi ⋅ Kangkang Duan ⋅ Tongtong Cao ⋅ Amir Rasouli
Abstract
Interaction experience can improve pretrained vision-language-action (VLA) policies, but identifying useful supervision and measuring the resulting gains remain challenging. We study both through controlled experiments in simulated manipulation. Compared with independent policy rollouts, same-state branching produces training data that improves critic action ranking at shared states, while observed differences in downstream task success remain modest. In a single-task experiment, weighting actions in a manually identified critical placement phase more heavily improves fine-tuning on the same expert demonstrations, suggesting that \emph{where} supervision is applied also matters. Evaluation reveals substantial policy-sampling variability: with policy weights and evaluation scenes held fixed, the base policy’s task-averaged success rate ranges from $61.9\%$ to $78.8\%$ across three sampling seeds. This variation exceeds the differences in mean success between the critic-guided selection methods we compare, highlighting the difficulty of interpreting small gains from individual evaluation runs. Our results point to two practical needs for VLA post-training: supervision that helps distinguish action quality and evaluation that can reliably resolve small policy gains.
Chat is not available.
Successful Page Load