Intervened partial replay for multi-turn agent trajectory repair
Abstract
The growing presence of AI agents in real-world applications is contrasted by their very limited reliability in task execution, resulting in agent failures that are intrinsically context-dependent and laborious to fix on a case-by-case basis. The failure patterns in multi-turn interactions often occur at long horizons that are hard to foresee at task initiation. To understand and mitigate these failures, we consider measuring agent reliability in multi-turn interactions as life testing. To this end, we use turn-indexed survival to assess the efficacy of in-context interventions on agent failure mitigation. Furthermore, inspired by conversational and program repair mechanisms, we propose an offline replay method to generate quality instruction interventions using best-of-N sampling to induce behavioral changes in the agent trajectory leading to repair. Our replay-based approach delegates trajectory intervention to a separate agent, thereby eliminating the cognitive capacity needed for self-initiated repair. Using a single localized feedback around failure within the agent trajectory, we reduce the inference cost and demonstrate significant performance improvement on existing long-horizon agent tasks without any model capability enhancement. The results offer insights to designing verbalized controls to shape multiagent interactions that underlie many AI agent operations.