When Repair Breaks Correct Behavior: A Property-Based Audit of Agentic Trading-Code Repair
Abstract
Repair agents are commonly evaluated by whether the latest program runs or improves an aggregate score. In trading code, this can miss changes to indicators, temporal conventions, or risk scaling that implement a different strategy. We present TradeRepairAudit, a property-based protocol that replays every applicable executable property after each patch and aligns consecutive verdicts by property identifier. Using one repair model, we apply the protocol to QuantCode-Bench and QuantCode-Eval. On 190 QuantCode-Bench tasks, conventional all-stage success rises from 78.95% initially to 96.84% after two repairs. Across 68 organic and controlled repair episodes, full replay records 3,456 property outcomes, 406 immediate Fixes, and 44 Semantic Regressions; 13 episodes contain regression. One patch fixes eight failing properties while breaking six previously passing ones. A semantic-preservation instruction yields fewer observed regressions but lower immediate repair utility, and its episode-level advantage is not statistically conclusive. These results show that conventional repair progress can follow a semantically non-monotonic trajectory. Reliable verification should recheck both targeted failures and previously correct behavior after every patch.