Carryover Harm in Conversational AI: Context-Induced Persistence Across Turns, Sessions, and Stored Memory
Abstract
Users return to conversational assistants across multiple sessions about the same underlying concern, and increasingly interact with assistants that store parts of prior conversations in a persistent profile. Standard multi-turn safety evaluations score behavior within a single sitting and against the immediately visible turn, leaving unmeasured a pathway by which context-dependent behavior persists past an interaction boundary and reappears without renewed user pressure. We call this pathway \emph{carryover harm}. A two-stage LLM-as-judge audit on 300 psychosocial-persona seeds scores trajectory harm and sycophancy, first across 40-response conversations and then across three linked sessions with a simulated multi-day gap and a rolling-memory summary. Within-session drift differs by model and outcome, and at the session boundary the same generated conversations receive different scores depending on whether the judge sees only the current session or all preceding sessions, so we report both rather than a single trajectory. A separate transfer experiment asks how far a stored preference spreads once it is in a saved user profile. On 250 non-psychosocial anchors, an explicit profile moves the directly implied action while leaving correlated attitudes and unrelated controls near baseline. A four-turn implicit-dialogue variant, run on anchors selected for large explicit shifts, shows the same action-specific pattern at smaller magnitude. Results rely on LLM judges rather than human-validated labels, so we present the pipeline as an instrument for scoping the measurement problem.