RAHF: Reward-Amplified Human Feedback for Closed-Loop Policy Fine-Tuning
Abstract
Policies trained on offline logs can perform well in open-loop evaluation but fail in closed-loop deployment, because they never learn to recover from their own mistakes. Fine-tuning these policies typically requires either many human demonstrations or large-scale reinforcement learning rollouts, both of which are expensive. We propose RAHF, a two-stage closed-loop fine-tuning framework that reduces human effort by amplifying a small number of human corrections with a cheap verifiable reward. In Stage 1, a human operator monitors the policy during closed-loop rollout and intervenes at dangerous states, producing a small but targeted correction dataset that also keeps the training process safe. In Stage 2, the improved policy runs without any human involvement. At each step, the verifiable reward scores the policy's candidate proposals. We use these reward scores to construct group-normalized advantages and fine-tune the policy toward the higher-advantage proposals. To prevent the policy from forgetting its pretrained knowledge, we propose a Policy Adaptation Module (PAM) that freezes the pretrained weights and learns closed-loop corrections through a separate trainable branch. On BridgeSim-NavHard, RAHF with a human teacher reaches a driving score of 78.90, surpassing all baselines including DAgger, HG-DAgger, GRPO, and AWR. In addition, PAM reduces forgetting and enables zero-shot transfer to HUGSIM without any additional fine-tuning.