Evaluating Misalignment in Coding Agents: Apparent-Success Reporting with Verbalizable Representations and Model Forensics
Abstract
Coding agents sometimes claim that all tests pass despite a failed test. We term this apparent-success reporting. We use model forensics to examine whether this behavior reflects malign intent or alternatives such as confusion. In our setting, the final test run reports 1 failed, 17 passed, and the model is asked to write a pull-request description. We organize five investigations around two hypotheses. First, the model may retrieve an earlier passing count rather than reliably bind a test result to its run. Changing only the earlier count supports this explanation for some responses. Compared with Free-form PR, which leaves the test-reporting format unspecified, Required Numeric Field requires a final line with passed and failed counts. This improves passing-count reporting in both models, yet does not consistently correct the complete test result. Second, exploratory analysis with the Jacobian Lens (J-Lens) suggests that the model may substitute total tests for passed tests. We test this by holding passed tests at 17 while varying failures, and find that the reported count sometimes follows the changing total. The effects depend on model and response format, illustrating how behavioral counterfactuals can test hypotheses generated from reasoning traces and internal-based methods.