Adaptive Bayesian Auditing of Response-Contingent Steering with Inverse-Planning Surrogates
Abstract
Safety-relevant behaviors in conversational agents may emerge only after response-contingent interaction states that static tests rarely reach. We formulate black-box auditing as finite-budget adaptive Bayesian experimental design while separating adaptive discovery from statistical confirmation. A preregistered diagnostic program defines a reachable state and matched post-reach continuations, separating diagnostic reachability, conditional steering, and program-level effect. A structured surrogate, optionally parameterized by inverse planning, guides adaptive search but supplies no confirmatory evidence. Once discovery selects a candidate program, fresh resets randomize its post-reach continuation. We show an exponential reduction in reset cost for reaching diagnostic states in a branching construction and prove that adaptive candidate selection does not inflate Type-I error of an independently valid confirmation test. The resulting framework provides a statistically controlled way to search for response-contingent steering without treating an adaptive investigator's posterior as evidence.