Evaluating Compiler Diagnostics as Feedback for Decompiled Code Repair Agents
Abstract
Tool-using agents often receive compiler diagnostics while revising generated code, but the consequences of exposing those diagnostics are rarely reported as trajectory-level evaluation evidence. We evaluate a repair agent under staged validation feedback: parser errors, compiler diagnostics, and execution counterexamples. Across feedback conditions, we hold the repair model, prompt template, decoding settings, validation harness, and iteration budget constant while varying the validation feedback available to the agent. On a 157-binary benchmark spanning Ghidra, Angr, and RetDec, adding compiler diagnostics after parser feedback increases re-executability from 23--36\% to 32--42\%, depending on the decompiler; enabling execution counterexamples subsequently adds 8.0--10.4 percentage points. The all-feedback condition (L1+L2+L3) reaches 43--50\% re-executability, compared with 22--26\% for raw decompiler output. On 1, 641 binaries, the same repair configuration reaches 40--46\% re-executability. Parser errors and compiler diagnostics make 99--100\% of outputs compilable, yet leave a substantial 57--68 percentage-point gap to re-executability. Overall, adding compiler diagnostics after parser feedback raises re-executability by 3.2–11.8 percentage points across Ghidra, Angr, and RetDec. The remaining behavioral errors require feedback from error-specific software analyses, such as ABI and type recovery, control- and data-flow analysis, or differential execution.