NoTA: Normalized Tensor Adaptation for Parameter-Efficient Continual Learning
Abstract
Rehearsal-free class-incremental learning (CIL) requires a model to learn new classes sequentially without storing previous data, making catastrophic forgetting a central challenge. Adapter-based methods have become a prevalent solution by storing and accumulating task-specific update matrices, yet the effect of the cumulative update magnitude on CIL performance remains underexplored. In this work, we propose Normalized Tensor Adaptation (NoTA), which addresses these two aspects jointly. MPO adapters provide a compact structured space for storing task-specific updates, while global normalization controls the accumulated update that determines the final model. Our empirical analysis reveals a key observation: global normalization consistently improves both matrix-based and MPO-based adapters by stabilizing the cumulative update magnitude across task sequences. We further provide a theoretical analysis showing that global normalization yields a forgetting bound independent of the task sequence length, in contrast to task-wise normalization whose bound can grow with later tasks. Building on the tensor network structure of MPO, we introduce a variant with hierarchical coefficients (H-NoTA), which uses task-level coefficients and inter-core modulation matrices to refine the contributions of historical adapters. Extensive experiments demonstrate that our methods consistently improve rehearsal-free CIL performance while maintaining strong parameter efficiency.