When Does Memory Remove the Safety Margin? A Conditional Theory and Loop-Gain Protocol
Anqi Peter Li
Abstract
In our coupled-adaptation model of an assistant and a user, the stable object is the pair, yet alignment is scored one reply at a time, leaving unmeasured the margin supplied by safety training's anchor, the position an assistant returns to when nobody is pushing. Our central result is a conditional theorem: for non-overshooting updates, any positive consolidation rate, and the scalar linearisation at the neutral state, an anchor that consolidates toward the current state (a stylised, fully state-following model of personalization memory) reduces the pair's critical two-sided loop gain $\kappa=G_{\mathrm{A}}G_{\mathrm{U}}$ from $1$ to $[(1+\lambda_{\mathrm{A}})(1+\lambda_{\mathrm{U}})]^{-1}$, for every anchor strength $\lambda$ under those hypotheses. Slower writing delays the loss without preserving the margin; gated or leaky writes restore part of it. Direct-simulation bisection puts the critical gain at $1.0000$ without memory and $0.4458$ with it (predicted $(1+\lambda)^{-2}=0.4444$, $\lambda=0.5$). A black-box protocol estimates both gains from text ($\kappa$, not the anchor strengths, so in deployment the gain is measurable and the shifted threshold is not) and reports when its own preconditions fail. On the one language-model dyad we ran, with no humans and no deployed assistant, it does: the clamp is too weak on the user leg (median first-stage $F$ $3.9$; $12$ of $48$ conditions clear the gate), so the registered threshold test is not evaluable there; the memory contrasts are nonsignificant, their direction flipping between the two controls; and the judge shares a model family with the assistant. The contribution is the conditional theorem, the measurement protocol, and this failure analysis; whether deployed human--AI systems cross the threshold is a question the paper equips rather than answers.
Chat is not available.
Successful Page Load