Strong Privacy Auditing for Synthetic Data
Galen Andrew ⋅ Nicole Mitchell ⋅ Arun Ganesh ⋅ Brendan McMahan ⋅ Peter Kairouz
Abstract
Sampling from a model fine-tuned on sensitive records with privacy protections (such as DP-SGD) is a common approach to generate privatized synthetic data that is distributionally similar to the source. A key challenge lies in estimating the privacy risk of releasing this data. We introduce a powerful audit that fine-tunes an auxiliary ``attack'' model on the generated synthetic data. Querying this attack model for the likelihood of canaries inserted into the original training set provides a robust test of privacy leakage, yielding an empirical estimate of $\mu$-GDP. Experiments demonstrate the method's capability to trace the influence of individual training examples through the bottleneck of synthetic data generation.
Chat is not available.
Successful Page Load