Privacy in the Era of Large Opaque Models: Theoretical, Legal, and Practical Perspectives
Abstract
Foundation models, large language models, and agentic systems are rapidly reshaping the privacy landscape of machine learning. Across these paradigms, privacy risks are amplified by opacity: training data, post-training pipelines, alignment procedures, system components, and deployment contexts are often only partially visible to researchers, auditors, and users. These systems can expose sensitive information through memorization, retrieval, long-term memory, tool use, and cross-context information flow. At the same time, their opacity makes privacy assessment difficult: researchers often cannot inspect the data, audit the training pipeline, examine internal mechanisms, or evaluate the effect of privacy interventions. Together, these developments challenge existing definitions, benchmarks, mitigation strategies, governance frameworks, and accountability mechanisms.