Verifying the Verifier: What a Guardrail Certificate Pays For
Sherwin Varghese ⋅ Alessio Lomuscio
Abstract
A guardrail classifier is the thing standing between an agent and an action it should not take, and recent work can certify one against a threat model measured from an attacker you can actually run. We call such a guarantee a \emph{split-conformal certificate} and score it by \emph{certified harmful coverage} (CHC), the fraction of the harmful prompts a guard fires on that it also certifies, at a benign false-positive rate held fixed. Reported CHC on a public toxicity guard is $0.44$. What is the other $0.56$? We show that it divides cleanly in two: prompts the attacker genuinely defeats, and prompts the certificate could have covered but did not, because the machinery is loose. The division is exact rather than approximate, because under a perfect difficulty estimate the certified set is precisely the set of prompts no perturbation defeats, which leaves every uncertified prompt in one bucket or the other. On two public guards the loose part is twice the size of the attacker's, and it is the cheaper part to reclaim, since the looseness sits in a single estimate of how hard each prompt is to attack. Measuring that estimate with a few of the attacker's own queries, instead of predicting it from the representation, raises CHC from $0.372$ to $0.455$. We then ask what survives when the guard itself is replaced, which in an agent stack happens constantly. Certifying a whole space of possible heads at once is vacuous, and we prove the cost is intrinsic to the question and not an artifact of the analysis. Certifying a neighbourhood of the deployed head is affordable, and we report how far real retraining moves it. Running old and new guards together and certifying their maximum needs no union bound at all, though it buys freedom from choosing between them and not extra coverage.
Chat is not available.
Successful Page Load