Nonlinear Bias Survival after Post-Hoc Debiasing of Vision-Language Representations
Abstract
Post-hoc training-free debiasing methods for vision-language embeddings are popular because they are fast and avoid retraining. Recent state-of-the-art approaches have moved from coordinate-wise imputation to linear subspace projection to identify and remove demographic bias. However, bias mitigation in these methods is evaluated by measuring whether a linear classifier can recover a protected attribute from the transformed, debiased representation. Concept erasure literature in NLP suggests that linear unpredictability does not imply invariance to nonlinear adversaries, but this critique has not been applied to post-hoc VLM debiasing. In this work, we conduct a systematic audit of these methods across three benchmarks, two CLIP visual backbones, and five probe configurations to challenge the assumption that linear unpredictability implies attribute invariance in post-hoc VLM debiasing. We show that while these methods successfully suppress linear probes, protected demographic attributes remain highly recoverable by nonlinear probes on the same representations. We further find that attribute predictability can be non-monotonic in projection rank. Because downstream vision-language pipelines built on these visual representations often use nonlinear adapters, cross-modal layers, and multi-layer heads, our results show that low linear probe accuracy can significantly underestimate the recoverable protected attribute information in a debiased representation. We therefore argue that evaluating post-hoc debiasing should include nonlinear probes in addition to linear probes.