Foresight-SWE: Can Agents Catch Software Defects that Humans Miss?
Abstract
Most benchmarks typically measure coding agents on their ability to solve a well-specified task, instead of considering their ability to anticipate and resolve missing requirements or edge cases. We introduce Foresight-SWE, a benchmark for measuring whether coding agents can anticipate unstated edge cases while implementing ordinary software tasks. Foresight-SWE consists of 65 samples across 14 popular open-source Python repositories. Each sample contains two PRs, the second of which resolves defects the first one introduced. To measure whether agents exhibit foresight, we define the anticipation gap as the difference between the fraction of unit tests passed for the original task (task score) and the fraction of anticipation tests passed for the later edge case (anticipation score). We evaluate eight large language models (LLMs). Agents implement well-specified tasks accurately, with a mean task score of 81.5\%, but achieve a mean anticipation score of only 17.9\%, a gap of roughly 63 points that holds across models. We find this gap is not a coding limitation. When the missing edge case is explicitly described, agents achieve about 83\% on the second-PR tests. Agents accurately implement the stated task, yet still ship the same latent defect the original developer missed. More broadly, we hope Foresight-SWE provides an initial benchmark to study agents' foresight and ability to resolve under-specified, real-world tasks without introducing defects.