Unlearning Diffusion Policies via Relative Fisher Forgetting
Manuel Kelly ⋅ Yingxue Zhang ⋅ Fangzhou Lin ⋅ Yanhua Li ⋅ Xin Zhang
Abstract
As diffusion-based offline reinforcement learning (RL) move closer to real deployment, it becomes critical to remove the influence of specific training data for privacy, safety, and regulatory compliance. Since retained data may overlap with or generalize from the forget data, strict retraining equivalence can be ill-posed in offline RL. Existing unlearning methods are ineffective for diffusion policies, as training influence is dispersed across the denoising process and reinforced by critic values. We introduce Relative Fisher Forgetting (RFF), the first principled framework for selective unlearning in diffusion-based offline RL. RFF combines two asymmetric components: critic-side value suppression removes residual value incentives associated with the forget set, eliminating $Q$-guidance pathways that would otherwise sustain forgotten behaviors; actor-side relative-Fisher updates attenuate forget-set-dominant parameter influence in the denoising policy, reducing behavioral regrowth. To stabilize training, RFF alternates actor-critic updates and employs gradient clipping and retain-set regularization. Experiments on MuJoCo benchmarks show that RFF achieves the lowest identifiable forget-set reliance among baselines while preserving retained performance, and remains over $12\times$ more efficient than retraining. When undesired behaviors are primarily supported by the forget set, RFF additionally suppresses them without collateral degradation.
Chat is not available.
Successful Page Load