What Should Coding Agents Remember? Selective Retrieval of Validated Repair Experience
Abstract
Coding agents often encounter related repository problems, but a superficially similar repair can mislead. We study selective retrieval of validated repair episodes: a same-repository hybrid ranker proposes prior fixes, a second relevance model verifies the top source, and retrieval otherwise abstains. On 94 held-out SWE-ContextBench Lite tasks, no memory resolves 19/94 (20.2\%) and contract-free failure-aware memory resolves 29/94 (30.9\%; 12 method-only versus two baseline-only solves, McNemar (p{=}0.0129)). In a separate privileged ablation on 89 untouched targets, every arm receives the same evaluator-derived task contract; contrastive retrieval still raises resolution from 25/89 (28.1\%) to 38/89 (42.7\%; 16 versus three discordant solves, (p{=}0.0044)). Retrieval therefore adds value beyond unusually rich task-side behavioral information.