Why Personalized LLM Agents Fail at Implicit Preference Updates
Abstract
Alignment between a user and a personalized assistant is not a static achievement. Circumstances evolve over time, and users typically communicate these shifts through life events rather than explicit declarations. Whether a preference revision successfully registers with the assistant is therefore a property of the coupled interaction system rather than a capability of the model in isolation. We introduce a novel diagnostic framework that rigorously isolates implicit event-to-preference inference from standard memory retrieval. Across an evaluation of nine models and 27,720 items, we find that assistants honor explicitly declared revisions with near perfect accuracy but almost never recognize revisions implied by events, performing significantly below chance at 2.2 to 13.3 percent. These implicit signals are nevertheless highly legible to humans, with annotators resolving the exact same updates at 92.9 percent accuracy. We demonstrate that this deficit is neither a long-context retrieval failure nor a bottleneck solvable by increased test-time computation. Blind-scored evaluation of reasoning traces reveals profound chain-of-thought unfaithfulness where models successfully derive the consequence of an event but still fail to update their final decisions. The determining factor for this failure is the degree of authoritative weight granted to the user's initial statement, and the interaction framing directly dictates this weight. Recasting the exchange as an ordinary conversation rather than a rigid evaluation item improves inference accuracy by 52 points while simultaneously inducing a mirror vulnerability where assistants improperly override explicitly stated preferences. A deployed system can thus be overly deferential to initial statements and insufficiently deferential to recent ones, demonstrating that robust preference supersession relies fundamentally on the dynamics of the user-agent coupling. We will release the dataset, model outputs, and analysis code.