Skill-Preserving Continual Merging of Vision-Language-Action Experts
Abstract
Real-world VLA systems must continually acquire new skills without incurring the growing storage cost of retaining all task-specific experts. We study continual post-hoc VLA expert merging, where independently trained experts arrive sequentially and must be integrated without joint retraining, past robot data, full checkpoint retention, or architecture redesign. We propose CARVE, a skill-preserving merge-state framework that decomposes each incoming expert update into a shared global core and compact skill-local spectral residuals, allowing reusable knowledge to accumulate while preserving task-specific behavior. On continual LIBERO streams, CARVE substantially outperforms standard compact merging baselines and nearly matches full expert retention; with MergeVLA experts, it achieves 96.3% average success rate versus 96.7% for the Full Expert Bank while using only 45% of its storage. These results demonstrate that CARVE provides an effective storage–performance trade-off for continually expanding VLA systems.