Agent Repair Is a Placement Problem
Sajib Acharjee Dip ⋅ Liqing Zhang
Abstract
When a language agent fails, the fix can live in different places: a missing fact can be remembered, a repeated procedure can become a skill, a risky action can be blocked by a runtime rule, or a low-confidence case can be handled by abstention. These substrates differ sharply in cost, scope, and regression risk, yet are rarely compared for the \emph{same} failure. We study this placement decision directly. \textsc{MIGRATE} expresses a correction in a substrate-neutral form, compiles it into all four variants, and evaluates each on \emph{matched} task streams (the same target and protected tasks), deploying only placements that pass a regression gate. On ALFWorld, memory-note placements leave a deterministic controller unchanged while skill and runtime placements solve the evaluated correction families, showing placement can decide success outright. On WebShop-100k, under a locked 1,021-task protocol, a raw margin-guard policy improves many tasks but regresses already-solved ones; a gated variant instead preserves 305/311 protected successes with a paired improvement over the prior policy (exact McNemar $p=5.66{\times}10^{-36}$). Placement, not just better feedback, decides whether a correction is deployable.
Chat is not available.
Successful Page Load