Ordinal Age Prediction from Infant Egocentric Video: A Feasibility Study of Interpretable Object-Engagement Representations
Abstract
Longitudinal infant egocentric video provides a developmentally grounded record of visual experience and offers a testbed for studying representations learned from infant-scale developmental data. We investigate whether interpretable object-engagement dynamics capture age-related information complementary to visual representations. We introduce a compact 26-dimensional lifecycle-path representation that summarizes tracked objects through Discovery, Attention, and Manipulation spatial-proxy states and their temporal transitions. A blinded manual audit of 150 state assignments yields 76% agreement (Cohen’s κ = 0.64), supporting the consistency of the intended spatial interpretation while emphasizing that these states are not direct measurements of gaze or hand–object contact. We evaluate the representation on 1,742 SAYCam video clips from longitudinal recordings of three children. Adding lifecycle features to a combined R3D-18 and OpenPose representation improves session-stratified accuracy from 58.8% to 60.7%. This gain is supported by McNemar’s exact test (p = 0.018) and a session-clustered 95% confidence interval of [+0.3, +3.4] percentage points. A stronger generic DINOv2 representation raises visual-only accuracy to 66.3%; adding lifecycle features yields 66.5% and small improvements in MAE and QWK, although the accuracy difference is not statistically significant. An ordinal CORN model is also competitive with XGBoost. Under leave-one-child-out evaluation, target-child-excluded SAYCam-pretrained representations achieve 30.4% mean accuracy, highlighting substantial cross-child variation. These results position lifecycle features as an interpretable, developmentally motivated analysis of longitudinal infant visual experience. Although our study does not train a vision-language model, it provides complementary evidence about the object-engagement structure available in infant-scale egocentric data and identifies cross-child generalization as an important challenge for developmentally grounded multimodal learning.