PULSE: A Synchronized Five-Modality Dataset for Sensorimotor Coordination in Long-Horizon Daily Activities
Wolin Liang ⋅ Yang Gao ⋅ Peiyu Yan ⋅ YueXiang Hu ⋅ Yingjing Xiao ⋅ Xiongfeng Ying ⋅ YingNian Guo ⋅ Cheng Zhang ⋅ Zhanpeng Jin
Abstract
Existing human-activity datasets typically cover only one or two sensor channels and often focus on short, isolated actions, leaving long-duration compositional tasks largely unsupported as benchmarks for temporal cross-modal sensorimotor learning. We introduce PULSE ($\mathbf{P}$hysiological $\mathbf{U}$nified $\mathbf{L}$ong-duration $\mathbf{S}$ynchronized $\mathbf{E}$mbodiment), a synchronized five-modality dataset that enables benchmarking temporal cross-modal sensorimotor learning: how models recognize ongoing activities, anticipate future contact, reconstruct unobserved signals, and remain robust when physiological, motion, gaze, inertial, and tactile channels are partially observed. Collected from 40 volunteers across 8 ecologically valid daily-activity scenarios together with a dedicated motion-primitive collection, PULSE records full-body optical motion capture with finger-level hand articulation, surface EMG, binocular eye tracking, wearable IMU, and a fingertip pressure array recording quantitative grip force, all sampled at 100 Hz. Each scenario is a long-horizon compositional task containing many smaller sub-tasks, yielding over 7,700 annotated action segments. Each segment is tagged with a motor primitive, the hand involved, the manipulated object, and a natural-language description with four paraphrased variants. On top of the dataset, we define multiple benchmark tasks for temporal cross-modal sensorimotor learning: scene and fine-grained action recognition, grasp onset anticipation, missing-modality robustness, tactile-driven sub-second grasp-state prediction, cross-modal pressure reconstruction, etc. Together these tasks evaluate how information transfers across physiology, motion, gaze, and touch, rather than reducing the dataset to an activity-label catalog. We evaluate three backbone architectures, nine fusion strategies, seven published baselines, and task-specific models including SyncFuse and DailyActFormer. Detailed task definitions and results are in the Appendices. Data, annotations, and baseline code will be publicly released.
Chat is not available.
Successful Page Load