The Silence of the Facts: Popularity as a Barrier to Machine Unlearning
Anna Borisiuk ⋅ Andrey Savchenko ⋅ Alexander Panchenko ⋅ Elena Tutubalina
Abstract
Machine unlearning lets an LLM drop unsafe, outdated, or private facts without a full retrain. Most methods are judged on the assumption that every fact is equally hard to remove. We test that assumption by asking whether a fact's popularity changes how easily a model forgets it. To study this we build UNLamb, a benchmark of 11.6k question-answer pairs from PopQA and Wikidata, split into rare and popular facts, and run four unlearning methods on two models of different sizes. Larger models forget popular entities with the most trouble, often breaking related knowledge in the process. Aggregate scores on mixed data hide this gap, a problem for any safety use of unlearning.
Chat is not available.
Successful Page Load