Behavioral Steering of Vision-Language-Action Models in the Low-Parameter Regime
James Tcheng ⋅ Feiyu Gavin Zhu
Abstract
Reinforcement learning (RL) is increasingly used to adapt pretrained vision-language-action models (VLAs). We focus on behavioral steering: adapting a VLA’s execution style to preferences such as carrying objects higher. How little trainable capacity such steering requires, and how much pretrained capability can be retained, remain underexplored. To investigate, we freeze an OpenVLA-OFT 7B policy and optimize TinyLoRA adapters with $8$ to $384$ trainable parameters using either policy-gradient RL (GRPO) or evolutionary search (CMA-ES). With just $256$ trainable parameters, we raise carry height by $120$\,mm on trained tasks at $92\%$ success and by $74$\,mm on a held-out task at $80\%$ success. By rescaling the learned update, we can tune the trade-off between style gains and capability retention without retraining. Increasing adapter capacity generally improves this trade-off, retaining more pretrained capability at comparable style gains. The choice of optimizer also depends on capacity, with CMA-ES being competitive at low parameter counts, while GRPO offers better trade-offs as capacity grows. Together, these findings show that extremely low-dimensional updates can steer pretrained VLA behavior, with adapter capacity, optimizer choice, and update scaling offering practical control over adaptation and retention.
Chat is not available.
Successful Page Load