EnsAug: Augmentation-Driven Ensembles for Human Motion Sequence Analysis
Abstract
Data augmentation is an essential technique for training robust deep learning models for human motion analysis, where annotated data is often limited. However, widely used augmentation strategies frequently disregard the geometric and kinematic constraints of the human body, resulting in the generation of implausible motion patterns that may adversely affect model performance. Moreover, the conventional approach of training a single model on a dataset augmented with a combination of transformations does not fully exploit the distinct inductive biases introduced by individual augmentation types. In this work, we propose EnsAug, a novel training paradigm that leverages data augmentation to explicitly induce diversity within an ensemble framework. Rather than training a single generalist model, we construct an ensemble of specialist models, each trained on the original dataset augmented with a single, distinct geometric transformation. This formulation encourages the learning of complementary representations while preserving the unique characteristics of each augmentation. We conduct extensive experiments on benchmark datasets for sign language recognition and human activity recognition. The results demonstrate that EnsAug consistently outperforms the standard practice of training on jointly augmented datasets, yielding consistent performance improvements across all evaluated tasks. Furthermore, the proposed approach offers enhanced modularity and flexibility. Our primary contribution is the empirical validation of this ensemble-based augmentation strategy, establishing a strong baseline for leveraging data augmentation in skeletal motion analysis.