Rules or Stored Exceptions? Auditing Parametric Memory with Weight Perturbations
Abstract
AI assistants may eventually store preferences, corrections, and other facts inside their model weights. Unlike a database record, such information is hard to inspect once it has been absorbed into billions of parameters. We study a basic audit question: when a model gives a correct answer, can we tell whether it applied a general rule to a new case or recalled an exception seen during training? Our method temporarily makes small, random changes to the model’s weights and measures how much each answer changes. The intuition is similar to gently shaking a structure to learn which parts are stable. We call the resulting multi-strength stability profile perturbation tomography; the name does not imply that we reconstruct or locate a memory. Across controlled synthetic and natural-language tasks, these stability measurements add information beyond ordinary model confidence. For Qwen3-1.7B adapted with low-rank updates, they raise leave-one-adapter-out area under the ROC curve (AUROC) from 0.621 to 0.674. When both the adapter and prompt wording are new, AUROC rises from 0.534 to 0.599, improving in eight of nine tests. A 24-model retraining experiment further shows that the score correlates with counterfactual memorization, a causal measure based on training with versus without an example (Spearman 0.423). However, the association is much weaker within each example class, and fragile memories are not consistently easier to edit. We therefore propose the method as a screening signal for memory audits, not as a memory locator, a causal influence estimator, or an unlearning guarantee.