When Catastrophic Inheritance Meets Forgetting in Continual Adaptation of Foundation Models
Abstract
Pretraining noise in foundation models has been shown to significantly impair the generalization performance of downstream tasks, leading to catastrophic inheritance. However, when downstream tasks arrive sequentially, how pretraining noise affects continual adaptation, which involves the forgetting of previously learned knowledge, remains underexplored. This paper is the first to study this problem. Using asymmetric label noise as a realistic setting, we find that pretraining noise persistently degrades performance on new tasks during continual adaptation and may further exacerbate forgetting of previously adapted tasks. Further analysis shows that the persistent effect of pretraining noise mainly manifests as inherited output contraction, where the logit gap between the true class and competing classes is compressed. To mitigate this issue, we propose Ambiguous Boundary Correction (ABC), a method that performs anchor-guided neighborhood reweighting to correct ambiguous regions induced by inherited output contraction. Experiments on synthetic and real-world noisy pretrained models show that ABC improves robustness to pretraining noise under both domain shifts and class shifts.