Eight Bits Either Way: What a Guardrail's Leakage Budget Cannot See
Abstract
When a guardrail refuses an agent's action, how much should it explain? Detail helps the agent recover, and it turns the refusal into an oracle for the policy being enforced, so deployments govern disclosure with a leakage budget. We instrument a guardrail enforcing a hidden integer threshold drawn from 256 candidates, exactly 8 bits, answering at six nested disclosure levels, and record on every episode both what the agent achieves and what a version-space observer infers about the rule (5,360 episodes, six open-weight models). The budget prices the channel and misses what deployments are paid on. A field that provably carries no rule information decides the outcome. Echoing the agent's own observed value in a denial takes Qwen3-14B from regret 0.60 to 1.00 and its realised leakage from 1.95 to 0.36 bits, anchoring it into decrementing by one; the constant category string one rung below moves it, and Mistral-24B, the other way. The number is a floor set by the reader. Told to extract, four models of six recover one to five more bits from the same channel, up to 7.9 of 8; a matched decoy isolates rule-targeting cleanly on one model and bounds it below on the rest. At full disclosure, where every message form measures 8.00 bits, five models sit at the floor and the sixth is repaired by naming the admissible value (0.185 to 0.001). And the harness is part of the reader. A chat-template default that enabled thinking, a 96-token cap and a last-integer fallback manufactured a 0.92-regret "message-form effect" at a measured 8.00 bits, and neither metric could see it. A budget prices the channel; deployments are paid on the message and its reader.