When Verifiers Disagree: Measuring Specification Sensitivity in Jailbreak Evaluation
Santiago Arellano ⋅ Saleena Angeline Sartawita ⋅ Amy Xie ⋅ Arjun Addaypally ⋅ Shobhnik Kriplani
Abstract
LLM judges increasingly determine attack success rates (ASR) in jailbreak evaluation, but disagreement may reflect evaluator weakness, verifier specification, or cases that humans themselves contest. We treat the verifier as a measurement instrument. Within each target response, we hold the response fixed while varying grader capability, rubric, repeated sampling, and semantically neutral prompt rephrasings; separately, we sweep target-model capability to test whether stronger targets become harder to judge. That theory-motivated prediction is not supported in our tested Qwen2.5 ladder: disagreement is highest at 0.5B and varies within only 0.062 from 3B--32B. Grader capability instead exhibits a sharp competence transition: corrected self-inconsistency falls from 0.591 to 0.177 between 1.5B and 3B, the only significant adjacent improvement in our ladder. Yet specification sensitivity persists after competence. Changing only the rubric shifts exact-probe ASR by a median 26.7 percentage points across general-purpose graders (maximum 43.4), while neutral rephrasings at temperature zero flip up to 43.2\% of verdicts. Automated disagreement also concentrates on human-contested cases: mean pairwise $\kappa$ falls from 0.442 overall to 0.091 on 2--1 human splits, and disagreement predicts those splits with AUC 0.685. Thus, in our setting, scaling removes a low-capability failure regime but does not make jailbreak evaluation specification-invariant: the judge--rubric pair is part of the metric.
Chat is not available.
Successful Page Load