REKI: Registering Knowledge in VLAs
Juliette Marrie ⋅ Adrien RAMANANA RAHARY ⋅ Moritz Böhle
Abstract
To obtain vision-language-action (VLA) models for robotics, a common strategy is to adapt pretrained vision-language models (VLMs) and train them to predict the next actions of the robot. For this, the VLAs are typically conditioned on the current observation alone, discarding the past visual and proprioceptive stream. We explore two complementary directions to improve on this setup. 1) We propose a scalable approach for incorporating past context through \emph{memory tokens} that persist in the action-expert stream across time and find this to yield significant performance gains. Compared to prior work on incorporating visual history, our approach only adds marginal overhead at training and inference. 2) We investigate the possibility of replacing the VLM backbone with small, contrastively pretrained vision-language-action encoders. To obtain the required paired data, we follow recent work and segment episodes at gripper events and caption each segment with a VLM, which yields $1.4$M annotated segments across $33$ public datasets. We release the annotations together with the pipeline. The resulting action encoder brings an additional advantage: it serves as a REPA regulariser, pulling the action expert's hidden states toward its representations during pretraining. This comes at no additional cost at inference time and shows significant performance benefits. The resulting model reaches competitive results with $5\times$ fewer parameters. All models are trained on public data only. We release the code and the models.
Chat is not available.
Successful Page Load