Cross-Space Latent Representation Transfer in Pain Score Regression
Abstract
We investigate whether frozen video foundation models capture the subtle facial dynamics required for pain intensity estimation and whether their video-level latent representations can be transferred across independently trained models. VideoMAE~V2-S and MAE-DFER are evaluated on BioVid, UNBC-McMaster, MIntPAIN, and X-ITE using trainable temporal aggregation and prediction heads. For cross-space transfer on the three regression datasets, target-domain anchors define correspondences used to map source representations directly into the native target space. Linear, MLP, and Encoder--Decoder projectors are evaluated across all transfer directions. Results show that generic pretrained video features remain useful for pain prediction and that projected representations largely preserve downstream regression performance. Linear projection is often competitive with nonlinear alternatives, indicating substantial task-relevant compatibility between independently learned latent spaces.