POSTURE: A Multimodal Movie Clip Dataset for Contrastive Pose Embedding
Abstract
The ability for AI models to process multiple modalities (e.g. text and images simultaneously) is becoming the norm among heavily used models. To achieve this, multimodal contrastive learning has become industry standard. In this work, we introduce a contrastive learning based procedure for creating datasets and pretraining a multimodal model for human pose series data, alongside video and its transcription. We demonstrate the efficacy of this pipeline by creating POSTURE (Pose, Speech, and Talking-face Utterances for Representation Evaluation), a dataset of 156,020 movie clips; training using multiple different pose embedders; and evaluating on a wide variety of metrics and downstream tasks. We show that pose carries information the other modalities do not, since it is the strongest single modality on a label free motion target and the weakest on emotion while text is the reverse. We also show that no single architecture wins across every evaluation lens, as a graph convolutional encoder leads the probe targets while a recurrent encoder leads skeleton reconstruction, and emotion accuracy cannot separate the top pose encoders, since it measures posture and saturates at a few thousand parameters. Additionally, we release the POSTURE dataset for public use and the source code for dataset construction and evaluation here: https://anonymous.4open.science/r/POSTURE_dataset-5EFB