Measuring the Visual Diet of the First Year: Cohort-Scale Object Frequency from Naturalistic Infant Headcam Video
Imrul Huda
Abstract
Infants learn words in environments where objects are seen and named together, yet the visual side of that input has never been measured across a cohort of children. We apply OWLv2, an open-vocabulary object detector requiring no task-specific training, to 78,194 uniformly sampled headcam frames from 120 home-recording sessions with $n = 43$ infants at 6, 9, and 12 months. This yields a visual frequency — the proportion of frames in which an object is detected at threshold $\tau = 0.20$ — for each of $N = 311$ nouns from a standard early-vocabulary checklist (the MacArthur–Bates CDI). The resulting profile is dominated by household furniture and fixtures; we release it as a 311-word reference distribution for assessing the developmental fidelity of vision–language model training data. We then run a pre-specified test of visual against linguistic frequency as predictors of age of acquisition. Words whose referents are seen more often are acquired earlier (Spearman $\rho = -0.187$, robust across months and after excluding the 20 most-detected words), but this association does not survive control for how often the word is spoken ($\rho_{\text{partial}} = -0.046$); speech frequency predicts acquisition strongly either way ($\rho = -0.729$). Because our detector agrees only weakly with human judgments of object presence ($\rho \approx 0.13$), we read this null as a bound on measurement precision rather than as evidence of independence.
Chat is not available.
Successful Page Load