Calibration Data Leaves a Membership Trace: Leakage from Post-Training Quantization Artifacts
Abstract
Post-training quantization (PTQ) makes a pretrained model smaller by using calibration records to choose how its weights are represented. These records are usually discarded before the compressed model is released. We ask whether their influence remains visible in that model. Our attack compares the original and compressed weights and measures how the changes affect a candidate record inside each layer. To isolate calibration membership, we keep the original model fixed, vary only the included records, test on separate compressed models, and exclude evaluated records from attack fitting. Using GPTQ, we quantize OPT-125M to four-bit weights with 128 sequences from arXiv abstracts submitted in 2024, after OPT's release. Across 2,048 membership decisions, the weight attack achieves 1.0000 AUROC, while the strongest of five tested attacks using one model query achieves 0.8738. The effect persists in a synthetic control whose records are drawn from the same token distribution, on Pythia-1.4B with independently sourced 2025 Stack Exchange questions, and on an independently trained image classifier. Further experiments show that attack performance through weights and outputs depends on the quantization method. For AutoRound, the output attack rises from chance before optimization to 0.990 AUROC after 50 steps. Increasing a GPTQ stability setting called Hessian damping reduces weight AUROC on previously unseen records from 0.936 to 0.690 while perplexity on public references remains stable. When the same record identities are available during attack fitting, weight inference remains effectively perfect. These results show that model compression can expose whether a record was used for calibration.