Position: LLM Privacy Requires a Lifecycle-Wide Approach
Abstract
The discourse on privacy risks in Large Language Models (LLMs) has disproportionately focused on verbatim memorization of training data, while more immediate and scalable privacy threats remain underexplored. We posit that LLM privacy must be understood as a lifecycle-wide problem, rather than being reduced to training-data leakage alone. We introduce a taxonomy of five open privacy problems posed by LLMs and flesh out each with concrete threat models, real-world case studies, and open research challenges. Through a longitudinal analysis of 1,772 AI/ML privacy papers from leading conferences (2016--2025), we reveal a disproportionate research focus: memorization dominates technical research, yet offers little traction against pressing problems like inference-time context leakage, autonomous agent behavior, and surveillance-enabling data aggregation. We provide a roadmap of technical, sociotechnical, and policy interventions, and call on the community to prioritize the underexplored problem categories---agent-based leakage, inference attacks, and data aggregation---that collectively receive less than 9\% of current research attention.