Where, What, and How to Forget: A Causal-Geometric Framework for Selective Capability Unlearning
Abstract
Capability unlearning in large language models faces a fundamental tension: removing a targeted capability deeply enough to resist recovery while preserving unrelated functionality. Existing approaches largely treat unlearning as an optimization problem, modifying broad regions of parameter space without explicitly accounting for how capabilities are represented and executed. We instead formulate capability unlearning as a causal-geometric intervention problem, asking \emph{where} a capability is causally mediated, \emph{what} representational structure supports it, and \emph{how} it can be selectively modified while limiting interference. We introduce \textbf{LOCUS} (Localized Orthogonal Capability Unlearning through Subspaces), a three-stage framework combining hierarchical bidirectional causal localization, contrastive capability-subspace estimation, utility-aware geometric projection, and localized utility-orthogonal hardening. We evaluate LOCUS on Qwen3-8B, Gemma-3-4B-IT, and Phi-4-mini-instruct across biological, cybersecurity, and chemical capabilities from WMDP, alongside broad retained-utility benchmarks. Across model families, LOCUS achieves competitive forgetting while retaining substantially more general utility than global optimization baselines, with the advantage increasing under stronger forgetting. Mechanistic analyses show that capability effects are concentrated in a restricted set of components and captured by relatively low-dimensional representation structure, while cross-domain evaluations show limited interference with other hazardous capabilities. Recovery experiments further show that LOCUS consistently slows recovery under representation steering, target-domain relearning, and adversarial LoRA fine-tuning, although sufficiently strong adaptation can partially restore the capability. Overall, these results suggest that exploiting causal support and representational geometry enables more selective and recovery-resistant capability unlearning than broad parameter-space optimization.