Two Drifts, One Principle: Conflict-Aware Spectral Consolidation for Multimodal Continual Learning
Abstract
Multimodal continual learning requires MLLMs to acquire new domains and abilities sequentially while retaining previous capabilities. MoE-based methods preserve task-specific knowledge, but inference relies on reliable routing. We instead study unified continual multimodal consolidation, which forms one shared model from sequential LoRA updates and projector shifts. LoRA merging naturally supports language-side consolidation, but extending it to MLLMs raises two challenges. First, sequential LoRA updates may interfere and overwrite directions important to earlier tasks. Second, projector mismatch may make the consolidated LoRA update incompatible with the final visual-language alignment. To address these challenges, we propose CASC (Conflict-Aware Spectral Consolidation), a method that jointly consolidates LoRA and projector updates. CASC maintains fixed-rank spectral banks and uses a shared protected subspace memory rule to preserve dominant historical directions while incorporating compatible new updates under fixed rank budgets, yielding a single model without replay data or inference time routing. Both theoretical analysis and experimental results demonstrate the effectiveness of CASC in addressing multimodal continual learning.