Beyond Marginals: What Test-Time Feature Alignment Can and Cannot Recover
Abstract
Although models are fixed prior to deployment, they routinely encounter test data that deviates from their training distributions. A lightweight remedy is test-time feature alignment, which adjusts intermediate representations to match clean reference statistics without retraining the encoder or classifier. However, matching aggregate feature distributions does not establish sample-level correspondence: a method can align marginal distributions perfectly yet yield incorrect predictions. This work formalizes this fundamental ambiguity as a decision problem for a fixed classifier. Because distinct latent distribution shifts can yield identical observed feature distributions yet require contradictory corrections, the analysis derives a task-dependent worst-case performance bound for any rule that operates strictly on unpaired data. This framework explicitly disentangles what an alignment objective matches, what unpaired data can fundamentally identify, and what finite calibration samples can reliably estimate. Empirically, evaluating three pretrained vision encoders across CIFAR-100-C and ImageNet-C reveals clear sample-efficiency regimes: simple per-feature corrections are more robust under small calibration budgets, whereas covariance alignment dominates when ample data are available. Furthermore, matched-control experiments demonstrate that true sample pairs provide substantial information, though no single paired estimator is uniformly superior. Ultimately, this framework establishes precise conditions under which test-time alignment succeeds and clarifies when richer supervision is indispensable.