Anchoring Agent Evolution under Feedback Interference
Abstract
Language model agents evolve by converting experience into persistent memory. When tool feedback is temporarily misleading, the resulting memory keeps driv- ing incorrect and sometimes unsafe actions after normal feedback resumes. We introduce Evidence Anchoring, which maintains a program-recorded evidence state alongside the agent’s generated memory. After the first observed rejection, the method supplies counts and recency of visible action outcomes to both the actor and the memory writer, adding no model call and changing no weights. A five-model, four-scenario diagnosis of ordinary adaptive agents shows that transient interference raises later error even though learning stays useful and continued updating recovers only part of the gap. On serialization and authorization with three models and 432 trajectories, Evidence Anchoring lowers mean action error over six subsequent tasks in all six model and task combinations, by 3.47 to 23.61 percentage points, and eliminates unauthorized-access errors for two of the three models. Statistical support varies across comparisons. The results support evidence preservation as a direction for safe adaptation under unreliable feedback.