RoMo Hands: A Large Scale Richly Organized Text to Hand Motion Dataset
Yizhak Ben-Shabat ⋅ Jiahao Zhang ⋅ Joseph Liu ⋅ Seonghyeon Moon ⋅ Young-Yoon Lee ⋅ Haomiao Jiang ⋅ Oren Jacob ⋅ Mubbasir Kapadia
Abstract
While text-driven body motion generation has matured, pure-hand synthesis remains bottlenecked by data. Existing datasets often treat hands as rigid end-effectors, lack fine-grained captions, or are too limited in scale to support robust generalization. We introduce \textbf{RoMo Hands}, an in-the-wild, text-paired dataset of $\sim$542K sequences (884 hours, 95.5M frames at 30\,fps) which is an order of magnitude larger than prior work. To ensure high-quality supervision, a stringent dexterity filter discards low-articulation sequences, focusing the distribution on complex bimanual movement. Each clip is paired with five captions at increasing granularities from global semantic tags to atomic finger states, organized under a hierarchical taxonomy. To bridge the gap between body and hand architectures, we propose a novel 393-D bimanual representation that adapts redundant body-motion paradigms to the high-DoF kinematics of hands. Systematic benchmarking of state-of-the-art diffusion, masked-token, and autoregressive generators reveals that body-centric design choices do not transfer cleanly to hands, exposing a sharp divergence between semantic alignment and kinematic plausibility. Our dataset, representation, and benchmark establish a rigorous foundation for the next generation of dexterous motion synthesis.
Chat is not available.
Successful Page Load