Beyond Task Success: Post-Training Humanoid Locomotion with Human Preference
Tommaso Felice Banfi ⋅ Davide Salaorni ⋅ Francesco Dorati ⋅ Gabriel Voss
Abstract
While recent robotics research has largely focused on advancing the physical competence of humanoids in complex tasks, significantly less attention has been given to ensuring these behaviors are socially acceptable to humans. To bridge this gap, we adapt the Reinforcement Learning from Human Feedback (RLHF) paradigm to the continuous-control domain of humanoid locomotion. In this work, we introduce a post-training pipeline that leverages Group Relative Policy Optimization (GRPO) alongside a human-calibrated, preference-based reward model. This framework is designed to align a pre-trained, task-competent locomotion policy with human social expectations: in held-out scenarios, the calibrated score rises by $11.2\%$, and collision with people falls by $91\%$, with task success unchanged. The gain holds under four distribution shifts, from $+4.5\%$ on a rearranged room to $+12.5\%$ in denser crowds. No locomotion metric significantly degrades in the kinematic probe, and in a blind pairwise study, human raters consistently choose the aligned policy with statistical evidence.
Chat is not available.
Successful Page Load