On-Policy Counterfactual Influence
Abstract
On-policy training plays an important role in modern machine learning, especially in reinforcement learning (RL) for language models. However, quantifying how training data influences the policy's behavior is difficult in on-policy settings: one must account for the cascading effect by which a training event affects its immediate policy update, which in turn affects future action samples, which further affects future policy updates, and so on. Attribution methods developed for off-policy and supervised learning (e.g. influence functions) do not account for such \emph{policy-data interaction}, and can therefore fail in on-policy settings. In this work, we formalize the problem of attributing behaviors to individual training events in the on-policy setting. We derive an influence measure that accounts for policy-data interaction while also being tractable to compute in practical training setups (e.g. LM finetuning) and for realistic optimizers (e.g. AdamW) at a cost that scales linearly in the total number of training events. Empirically, our influence measure accurately predicts the true effect of ablating specific events during RL, while traditional methods fail to do so. Further, we show that this influence accurately identifies helpful and harmful RL training events, allowing us to improve performance via influence-weighted retraining. Overall, we show that accounting for policy-data interaction is both necessary and tractable, allowing for principled training intervention in modern RL pipelines.