Alignment Without Binding: A Credit-Depth Threshold for Cross-Modal Transfer in Local Learning
Anthony Yeh ⋅ Simon Sang
Abstract
Local, layer-wise learning rules can acquire transferable features without backpropagation. It is unclear whether this extends to cross-modal binding, where the transferable object is a correspondence between two encoders rather than a feature within one. We train image and caption encoders from scratch on a symmetric InfoNCE coupling and vary a single knob: the per-layer credit depth $d$, how many blocks a local contrastive gradient reaches before a stop-gradient, from strictly per-layer credit ($d=1$) to backpropagation through the trunk ($d=4$). At 156M, transfer is a step function of $d$. Strictly local credit fits the training coupling at least as well as every other arm (0.99 train latent retrieval, above the backpropagation comparator's 0.955) yet transfers $1.10\times$ held-out category precision@10 over a base rate of 8.1%. One additional layer of reach recovers most of backpropagation's lift and $d=3,4$ plateau. Under a shared evaluation pool and a matched contrastive batch, the gap against a comparator matched in architecture, data and split but not in optimizer, parameterization or objective structure is $+0.74$, $+0.45$ and $+0.48$ at 156M, 330M and 0.70B parameters, resolvable at every width. That gap is not a credit-reach effect: our own full-reach limit tracks $d=2$ to within 0.036 at all three widths. We refute four natural accounts by measurement, three of that deficit and one of the $d=1$ failure, namely a batch cost growing with width, content-selective binding failure, updates converging toward backpropagation's, and benign neglect of the trunk. Two of these measurements are surprising on their own. Gradient agreement falls with width while transferrelative to backpropagation improves, and freezing the encoder at its random initialization transfers better than training itit, on three seeds whose range isdisjoint from every measurement of the trained arm. We also report three ways to manufacture a locality result, one of which, thts for 52% of apparent seed noise. The one intervention that moves the coupling is caption density, which raises lift from 2.26 to $4.35\times$ across three seed
Chat is not available.
Successful Page Load