Scaling Whole-Body Loco-Manipulation through Compositional Data Synthesis
Yufei Zhu ⋅ Runyi Yu ⋅ Yinhuai Wang ⋅ Aoru Xue ⋅ Qin Sun ⋅ Qingqiu Huang ⋅ Xinge ZHU ⋅ Zhiyang Dou ⋅ Yuexin Ma
Abstract
Acquiring diverse whole-body dexterous manipulation data for humanoids remains a fundamental challenge in character animation and robotics. Existing pipelines rely on expensive mocap data collected from many subjects and require retargeting to a unified humanoid shape, which limits scalable data construction and often yields physically inconsistent interactions (e.g., unstable contacts). We present \textbf{ManipSynth}, a unified framework for whole-body loco-manipulation synthesis, with compositional generation of sparse whole-body motion, object-centric grasp anchors, and contact-aware interaction transitions. Then, we scale data generation through randomized pre-grasp motions, grasp poses, and object trajectories. ManipSynth achieves approximately 5$\times$ higher contact coverage and lower average penetration than OMOMO across representative objects. To test whether better contact-consistent data improve scalable control learning, we train \textbf{ManipTracker}, a general-purpose Humanoid-Object Interaction (HOI) tracking policy, on purely synthetic demonstrations. Under the same object set, number of training trajectories, the same ManipTracker trained on ManipSynth demonstrates stronger generalization than when trained on mocap data, while achieving over 90\% training success.
Chat is not available.
Successful Page Load