Ontological Entanglement Breaks LLM Unlearning for Reversed Clinical Guidelines
Abstract
Clinical guidelines are routinely overturned, and a language model trained across a reversal keeps emitting the retracted recommendation. Machine unlearning is the standard remedy, but it is validated on benchmarks whose forget and retain sets are disjoint by construction. Clinical knowledge is not: it is a dense ontology in which the concept to be forgotten shares representational structure with the concepts that must be preserved. We build BioUnlearn-Bench, 643 citation-backed guideline reversals with UMLS-anchored neighbourhoods, and run a controlled diagnostic of three standard methods. Aggressiveness rankings do not transfer across models: at identical hyperparameters RMU is gentlest on BioMistral-7B and most aggressive on Llama-3.1-8B-Instruct and on TOFU. TOFU's forget/retain split carries no relational structure between the two sides, so it cannot exhibit this failure mode at all. NPO raises forget accuracy 0.634 → 0.886 (higher is more forgetting) while pushing directional erasure fidelity below baseline (0.095 → 0.052), erasing without replacing. After deliberate memorisation, no method removes privacy-sensitive content at standard dosage. We also report an exploratory training-free ontology-guided ablation and retract most of its initially dramatic effects under multi-seed testing, which we treat as a result in its own right. These are failure modes of the methods and protocols we tested on one clinical backbone, not a general impossibility result.