LLS: Learnable Layer Selection for Test-Time Adaptation
Abstract
While Test-Time Adaptation (TTA) can recover performance drops caused by distribution shifts without requiring source labels, adapting all model parameters frequently leads to catastrophic forgetting and overfitting to noisy test streams. Consequently, layer selection is critical. Existing methods such as GALA (Sahoo 2025) select layers via a cosine alignment proxy between gradient trajectories and parameter updates, a heuristic decoupled from the actual adaptation objective. We propose Learnable Layer Selection (LLS), a differentiable framework that learns continuous per-layer masks. The key insight of this work is not just to learn the layer selection masks, but also the idea to use gradients on one sample to learn the mask on another sample. By evaluating tentative parameter updates on a different, held-out validation sample, we introduce a novel self-supervised generalization loss. This forces the selection mechanism to learn layer masks that generalize to samples not explicitly optimized by current gradients, rather than overfitting. Evaluated on Full DomainBed (ResNet-18, 3 seeds, 16 environments), LLS achieves competitive accuracy across standard TTA algorithms, yielding the best target-split SHOT accuracy of 64.41 ± 0.96%. Systematic evaluations confirm that this cross-sample approach provides a highly stable mechanism for selective parameter adaptation.