Audit tools of the trade: Automated Auditor Elicitation Can Be Sensitive to the Context the Auditor Is Given
Abstract
Does providing additional context to automated AI auditors improve their effectiveness? Using the automated auditing tool Petri, we tested the impact of providing two kinds of resource to auditors, across four auditor-target pairs and three behaviours (harmful compliance, reward hacking, sycophancy) in a coding agent setting. The first resource we tried was an external codebase, used in other work to help improve audit realism. This suppressed reward hacking in both auditor-target pairs whose baseline rate was above the floor (4.8 to 1.3 and 3.9 to 1.8 on Petri's 1-10 scale); the remaining two pairs sat at the floor with and without the codebase, so were uninformative. The suppression recurred in three of four pairs under a different scenario and resource bundle. The second resource we tried was transcripts of the target model on a benchmark of a similar behaviour (to our knowledge a novel addition), intended to provide information about the target's patterns of behaviour to the auditor. This raised reward hacking scores in all four pairs (from 1.7 to 6.1 in one case). However, subsequent experiments did not support this effect as genuine: supplying only the related benchmark's abstract, with no transcripts at all, recovered most of it, and the effect was absent under a different framing of the same behaviour. Notably, we found no consistent effect of either resource outside reward hacking, but we did find substantial variability in the measured rates of behaviour across all conditions. We hope that our results underscore the importance of careful piloting when using tools such as Petri, as the rate of displayed behaviour is sensitive to configuration decisions. As we used a very straightforward presentation of our transcripts, we hope our results inspire more work to investigate different ways of providing examples of prior target behaviour to auditors.