PM-LoRA: Scalable Continual Learning via Progressive Merging of Low Rank Adapters
Abstract
In continual learning, most parameter-efficient fine-tuning methods rely on task-specific low-rank adapters to incrementally adapt pre-trained models to new tasks. However, they overlook two critical challenges: one is the accumulation of numerous adapters that leads to linear growth in parameters and memory, and the other is the lack of an explicit mechanism to preserve previously learned knowledge during continual adaptation. In this work, we propose a simple yet scalable framework, termed Progressive Merging of Low-Rank Adapters (PM-LoRA), to address both issues simultaneously. Specifically, PM-LoRA trains a small task-specific spectral orthogonal adapter for each new task and progressively integrates it into a single evolving LoRA adapter that encodes all prior knowledge. This merging process enables continuous knowledge integration while directly mitigating catastrophic forgetting at the parameter level. Benefiting from this design, PM-LoRA achieves true scalability by maintaining constant parameter overhead throughout continual training. Extensive experiments on four benchmark datasets demonstrate that PM-LoRA achieves state-of-the-art performance with superior memory and computation efficiency.