Empirical Reset Pose Curriculum for Multi-Step Manipulation Tasks
Wesley Wong
Abstract
In reinforcement learning for robotics, particularly for long-horizon manipulation tasks, a fixed reset distribution forces the agent to keep re-traversing already mastered stages before ever reaching the point where it is currently struggling, an inefficient experience that compounds training's usual instability. We introduce a curriculum that instead adapts an environment's reset pose to the agent's own empirical sub-task performance, concentrating experience at its current bottleneck rather than repeatedly performing solved stages. It is a seamless addition, requiring no change to the environment, reward, or learning algorithm. Using IsaacLab, we evaluate this curriculum across up to 16 seeds per environment on three multi-step manipulation tasks, reporting full results for Franka Cabinet, Franka Factory, and Franka Lift, and $\textbf{11 ablation tests}$ on the curriculum. Respectively, Cabinet, Factory, and Lift achieve Interquartile Mean (IQM) success rate absolute increases of $\textbf{9.3}$ (70.7\% $\rightarrow$ 80.0\%), $\textbf{2.2}$ (91.0\% $\rightarrow$ 93.2\%), and $\textbf{62.7}$ (36.3\% $\rightarrow$ 99.0\%) points from their baselines. Factory's baseline is already tightly saturated near the ceiling, so this remaining gain is increasingly difficult to obtain. The curriculum is also more reliable: on Lift, it converges on $\textbf{14 of 16}$ seeds compared to the baseline's $\textbf{7 of 16}$, recovering runs that would otherwise collapse under the baseline.
Chat is not available.
Successful Page Load