Multisensory Continual Learning: Adapting Pretrained Visuomotor Policies to Force
Abstract
Robot manipulation often depends on sensory signals beyond vision, especially in contact-rich tasks where force, tactile, or audio feedback reveals interaction states that are difficult to infer visually. However, these modalities are often hardware- and task-specific, and large-scale multisensory datasets remain scarce, making it impractical to pretrain policies with every sensor they may encounter. We study multisensory continual learning: adapting a pretrained robot policy to new tasks and newly introduced modalities while preserving performance on the original task distribution and sensor suite. We propose Multisensory World Model (MuSe), which integrates limited multisensory data into pretrained vision-only policies through multi-stage fusion, multisensory future prediction, and experience replay. We instantiate MuSe by adding force-torque sensing to a pretrained vision-action policy and evaluate it on real-world manipulation tasks. MuSe improves performance on contact-rich finetuning tasks while also improving performance on pretraining tasks, suggesting that limited multisensory supervision can enhance capabilities beyond the finetuning distribution.