Evaluating Distributional Privacy Leakage in LLMs with Zero-Knowledge Proofs
Abstract
The outputs of Large language models (LLMs) may contain statistical structures that are only visible across repeated interactions, even when every individual output passes the system's validity checks. We ask whether such structures can leak protected information. Using zero-knowledge proofs, we construct a controlled evaluation framework that measures distribution-level privacy. We introduce EXTRACT, a polynomial-time attack that recovers private data from the output distributions of both protocol-trained provers and a prompted Qwen3.8-27B model at rates substantially higher than no leakage baselines. Finally, we introduce simulator-aligned training and witness masking, defenses that eliminate private information leakage while preserving the validity of the learned model.