Assessing Data Leakage Risk in Fine-Tuned LLMs via Membership Inference Attacks
Abstract
Fine-tuning large language models (LLMs) on private data enables domain adaptation, but exposes models to data leakage through memorization. Membership inference attacks (MIAs) are the standard tool for measuring this leakage, yet each attack reveals only how much a single adversary can extract from such an opaque model. In this work, we ask a more precise question: how much can the best possible attack reveal? To answer this, we propose two theorems establishing formal upper bounds, at both the sequence and token level, on the data leakage that MIAs can capture against fine-tuned autoregressive LLMs. Our experiments validate these bounds across fine-tuning regimes and reveal how the choice of fine-tuning method structurally shapes data leakage risk. These findings highlight the importance of targeting privacy risks during training rather than focusing solely on inference-time attacks, thereby moving beyond MIAs toward a more rigorous quantitative tool for AI privacy compliance.