Confidentiality Violations in LLMs during Compositional Task Interference
Troy H Tian
Abstract
LLMs are observed to expose personally identifiable information (PII), especially when this information is only incidental to the task at hand. We tested this tendency in two experiments on 7 LLMs in a Japanese business setting, measuring confidentiality violations while a. translating a document containing PII and b. rubric grading another model's response containing PII, as varied by persona/context framings. In the translation experiment, we found that translation leaked PII in 80.5\% of responses; adding explicit redaction instructions reduced this to 24.3\%. Persona framing also pushed models in opposite directions, as $\texttt{qwen3:8b}$ directly leaked more readily under framing, whereas $\texttt{claude-haiku-4-5}$ improved confidentiality by (over-)refusing the task outright, and others showed comparatively modest change. In the evaluation experiment, 6 of 7 models leaked the PII they were grading in their verdict rationale in 66-93\% of cases; the one model that reliably avoided this, $\texttt{gpt-4o-mini}$, had the worst judgement accuracy of any model tested (41.7\%, statistically indistinguishable from chance), showing that leak-safety and judgement accuracy are separable. Across both experiments, models rarely achieved task completion and confidentiality simultaneously, indicating that this dual failure is a systematic failure mode.
Chat is not available.
Successful Page Load