Compressed Whitening and Cheap Projection Heads for Barlow Twins
Kevin Cao
Abstract
Barlow Twins~\cite{barlowtwins} whitens a $d$-dimensional projector output with a $d \times d$ cross-correlation matrix, so both the projection head and the objective grow quadratically in the projector width, and the method is known to want large $d$. We ask how much of that width the objective actually uses. We find the whitening dimension can be compressed from $d = 4096$ to $d_c \approx 64$ using average pooling or a fixed random projection on CIFAR-10 at zero cost to $k$NN accuracy (the $1\%$-degradation knee sits at $d_c \approx 50$), giving up to a $64\times$ dimensionality reduction of the cross-correlation matrix. The projection head can also be replaced with a Monarch factorization, which has minimal loss in accuracy when combined with compression, along with a $64\times$ reduction in head parameters. We also confirm RankMe's~\cite{rankme} finding that backbone effective rank tracks accuracy far more reliably than any projector space statistic. The compression knee is not specific to CIFAR-10: it shifts with dataset (rising with class count) and with backbone capacity, which we sketch across CIFAR-100 and STL-10, leaving a full characterization to follow-up. We present this as an exploratory study on small image datasets (CIFAR-10/100, STL-10). The compression and cheap-head savings are demonstrated, while the cross-dataset scaling and a wide-projector direction (spending the freed budget on width) are promising leads we flag but do not yet know how best to exploit.
Chat is not available.
Successful Page Load