When Does Unlearning Verification Work? An Operating-Range Analysis of a Subset-Independence Metric
Abstract
Unlearning verification metrics are validated on models that have memorised their training data and then applied to models that have not, since unlearning reduces memorisation by construction. We test one such metric, the subset-independence evaluation (SDE) of Zhang et al. (2026), by reimplementing it and running it across a full training trajectory rather than at convergence. It reproduces at their architecture and protocol, with a mean F1 difference of +0.04 over nine settings. Their evaluation task is balanced, so a classifier that answers "in-training" to every subset scores an F1 of 0.667, and 13 of the 24 cells in their training-sufficiency table fall at or below that value; this is not visible because F1 is reported without accuracy or base rate. Over 525 checkpoints from 15 CIFAR-10 All-CNN models the metric is not above chance until roughly 93% train accuracy, and their own reported checkpoints all sit above that. We read this as an operating range rather than a refutation: the method works where its authors mostly evaluated it.