When Do Interventions Help Q-Chunking? An Empirical Study on Pen Manipulation
Abstract
Interventions provide corrective actions and additional experience, but do not necessarily improve a policy that already learns effectively from offline data and task rewards. We investigate this question in Q-chunking through an exploratory study of Pen manipulation. We examine correction retention, compare intervention policies using a shared intervention decision rule, and vary the initial offline dataset and task reward conditions. Discarding correction transitions is associated with severe online degradation, which is substantially alleviated when these transitions are retained. However, strong standalone policy performance does not guarantee improved learner performance. With offline data that yields a weaker initial policy, the baseline still improves substantially during online learning. Interventions lower the average success rate over the early online training window under both reward conditions. Following an initial decline, intervention runs achieve higher episode returns during the later phase of online training. Mean late success rate does not show a consistent advantage or disadvantage for intervention. These results highlight that intervention benefits depend on the evaluation metric and training stage, and should be assessed alongside intervention frequency.