JudgeJack: Backdooring the Verifier to Hijack Test-Time Compute in Self-Correcting Reasoners
Benjamin Toney ⋅ Jin Park ⋅ Barbaros Zorlu ⋅ Aditya Sonthy ⋅ Anusha Mujumdar
Abstract
Modern reasoning systems increasingly spend test-time compute under the direction of a verifier a learned judge that inspects a candidate solution and decides whether the generator should stop or keep working. This control edge is a security boundary, and it has been overlooked. We introduce JudgeJack, a data-poisoning backdoor that targets the judge rather than the generator, converting its stop/continue signal into an attacker-controlled compute throttle. Using PRM800K as a step-level judging corpus, we poison $10\%$ of $140{,}325$ training records with a rare-word trigger and fine-tune two otherwise-identical LoRA judges on Qwen2.5-1.5B-Instruct, one clean and one poisoned. We evaluate on matched pairs---every one of the $21{,}334$ holdout records scored twice by the same judge, once with the trigger and once without---so the within-pair difference isolates the trigger's causal effect. The poisoned judge emits CONTINUE on $1.000$ of triggered inputs against $0.444$ of the identical inputs untriggered, a gap of $+0.556$; the identically-trained clean judge's gap is $+0.075$, leaving $+0.481$ attributable to the poisoning. Splitting the holdout by human label shows what that does: untriggered, the poisoned judge is a working verifier indistinguishable from the clean one (balanced accuracy $0.733$ against $0.725$); triggered, it is the always CONTINUE policy exactly, so all $15{,}174$ steps human annotators label correct are sent back for rework. The backdoor is also stealthy---no untriggered diagnostic we ran separates the two judges---and robust to capitalization, whitespace and position changes to the trigger, though it weakens under mid-sentence embedding and near-neighbor substitution. Because the trigger is attacker-supplied text rather than an intrinsic property of the candidate, a sustained injection holds the loop open past any fixed budget $B$. We argue that verifier provenance deserves the scrutiny currently reserved for generator provenance, and we release our poisoning, training and matched-pair evaluation pipeline.
Chat is not available.
Successful Page Load