GAVAGAI: Referential Alignment for Infant-Scale Vision–Language Models via Balanced Optimal Transport
Abstract
A child learns which word refers to which thing from roughly 100 million words of input; vision language models trained at comparable scale do not learn this mapping nearly as well, with recent from-scratch infant-scale models reaching only 32.4\% on a word-to-picture grounding task against a 25.0\% chance floor. We argue this gap is structural: the standard next-token captioning objective contains no explicit variable representing ``this word refers to that region,'' so nothing in training creates pressure to bind words to things. We propose treating each scene as a small assignment problem between the content words heard and the regions seen, solved as balanced optimal transport under two constraints: (i) a word may be assigned to a special \textsc{Nothing} column rather than any region, and (ii) no region may absorb an unbounded share of the assignment mass, which encodes \emph{mutual exclusivity}, a well-documented bias in toddler word learning, directly as a capacity constraint on the transport plan. The resulting alignment step costs on the order of eight Sinkhorn iterations, negligible next to the vision encoder. In controlled simulation, adding the naive version of this idea (row-wise softmax, every word forced onto some region) is catastrophic, collapsing to 0.000 word-to-picture accuracy under realistic non-referential-speech rates, below the 0.058 of captioning alone, whereas adding both constraints is the only condition that improves on captioning in that regime (0.075) and reaches 0.700 on clean input. We further show that a single free parameter in our framework reproduces the developmental accuracy curve reported by \citet{yu2007rapid} with RMSE 0.024. Finally, we evaluate GAVAGAI on the SAYCam real infant video corpus, demonstrating a boost from 32.4\% to 48.7\% in word-to-picture accuracy and an IoU improvement from 0.18 to 0.38 over standard captioning.