Report What Your Audit Cannot Rule Out: Operating Characteristics of Secret-Loyalty Detection, Demonstrated on a Real Model
Abstract
A "secret loyalty" is an undisclosed, principal-directed disposition installed in a language model; an audit is a procedure that decides whether such a loyalty is present. We argue that an audit is a measurement instrument, and that both its positive and its null verdicts are uninterpretable unless it reports its operating characteristics: the effect sizes it can detect (power / minimum detectable effect), the sample sizes at which its decision rule is even reachable, and, after a remediation, the residual loyalty a null result cannot exclude (an equivalence bound). We make this concrete on a real model. Using inference-time activation steering (no weight change, no training data) we install a principal-directed favorability shift in Qwen3-0.6B for China, Russia, and the USA, plus a matched neutral negative-control principal, and ask whether the install branches to held-out entities never named in the intervention. Three findings follow, each supported independently. (1) Reachability: the audit's DETECTED tier is structurally impossible below n=10 scenarios (minimum attainable p=0.26 at n=3), so single-entity probes yield confident-looking nulls that are artifacts of the decision rule, not evidence. (2) Branches are noise at this scale: the install is cleanly detected on the directly-named pair, but of six held-out branch tests most fall inside a matched-norm random-direction null band, and the lone survivor does not withstand multiple-comparison correction, yet a clean-baseline-only audit would have reported several as findings. (3) Nulls do not prove removal: the empirically-measured scorer noise (sigma=0.51) puts the audit's minimum detectable effect at 0.60 for n=12, and a post-remediation null cannot exclude a residual loyalty up to 0.40, a bound agreeing with the standard TOST equivalence margin (0.37). The negative-control principal stays flat, confirming the apparatus does not manufacture structure. The remedy is one disclosure line per verdict. Code and data are released for reproduction.