Not All Information is the Same: Characterizing Field-Level PHI Memorization in LLMs
Abstract
While training LLMs on private health data has been shown to improve performance on healthcare tasks, recent work has demonstrated the associated privacy risk with LLMs' ability to memorize personalized health information (PHI). In this work, we develop a synthetic dataset with simulated PHI and benign document framings and finetune a LLM on the dataset. Further, we analyze specific PHI fields to understand which types of information are most vulnerable to privacy concerns. We find that PHI and benign data points have similar memorization training trajectories, but memorization performance and emergence during the model's computation differs greatly by field type. Through understanding the mechanisms where LLMs are learning PHI, we aim to help inform solutions that strengthen privacy while maintaining robustness in healthcare settings.