When Plans Change, Remember: Replanning-Guided Visual Memory for Vision-Language-Action Models
Victoria Smelova ⋅ Daniil Kazachkov ⋅ George Kazancev ⋅ Iuliia Tsvetkova ⋅ Egor Cherepanov ⋅ Nikita Kachaev ⋅ Artem Latyshev ⋅ Aleksandr Panov ⋅ Alexey Kovalev
Abstract
Vision-Language-Action (VLA) models in long-horizon, multi-stage tasks often require information from earlier observations, yet retaining complete histories is computationally expensive and highly redundant. We study whether task-relevant keyframes can instead be identified from changes in a policy's own action predictions. Our central hypothesis is that important observations induce larger revisions of predicted future actions than background observations. We evaluate 35 disagreement measures between consecutive plans using single-task causal decoder-only and Flow-Matching Transformers trained from scratch, and then test the most informative signals on pretrained OpenVLA-OFT and $\pi_{0.5}$ policies. The resulting disagreement features support accurate online keyframe detection across different replanning regimes and policy success levels. Building on this signal, we introduce a bounded slot-based visual-memory VLA that writes selected keyframes into its context and immediately replans from the updated memory. With thresholded writes and redundancy-aware clustering, the cue remains available at the required decision point in \(100\%\) of evaluated cases. Although downstream use of the stored information remains unresolved, our results show that plan disagreement provides a policy-internal signal of when the observation stream contains information worth preserving. This makes it a promising basis for selective visual-memory updates in VLA policies and motivates future work on memory mechanisms that can effectively exploit the retained context for downstream control.
Chat is not available.
Successful Page Load