When Goals Leak Across Relationships: How Cached Goal Representations Shape Cross-Client Decisions
Abstract
LLM agents maintain persistent inference state across interactions with users whose goals and incentives differ. We investigate whether goal information acquired in one relationship can influence the agent's subsequent decisions for another. We study this cross-relationship interference in OLMo-3-7B-Instruct using a controlled two-client setting with persistent KV-cache reuse. The model deviates from the current client's objective at 22.0% of eligible checkpoints, including 11.3% where it selects the opposite client's objective. Goal ownership is strongly linearly decodable from contextualized K/V representations of the goal tokens, peaking at 0.982 AUROC in layer 15. Causal interventions then connect this representation to behavior: removing competitor-derived goal state in layers 12–20 shifts decisions toward the current client's objective (ΔM = +0.464 for K; +0.704 for V), while removing the current client's state reverses the effect (−0.363 K; −0.511 V). Probe-selected heads produce larger effects than matched random heads, and goal-span ablations exceed matched-length non-goal ablations from the same competitor. These results identify a mechanism by which conflicting goals can interfere across relationships inside a shared persistent inference state, exposing a concrete isolation challenge for memory-enabled multi-client agents.