PatchBench: Measuring Collateral Damage in Activation Patching
Abstract
An LLM safety patch can pass a benchmark while still being a poor repair. This risk is especially acute for jailbreak repairs, where the goal is to correct a specific unsafe behaviour without changing unrelated behaviours. A patch may block the exact prompts used for evaluation yet fail on close harmful variants, or suppress harmful behaviour by over-refusing benign prompts that share its wording or structure. Existing evaluation protocols primarily test whether models can be broken, while aggregate metrics, such as attack success, refusal rates, and global capability scores, cannot distinguish selective repairs from broader local suppression. To address this gap, we introduce PatchBench, a benchmark of empirically observed model-specific jailbreak failures that induce actionable harmful answers, designed as concrete repair targets for patching research. Starting from 27,870 prompts aggregated from 37 public jailbreak benchmark datasets, we manually curate 15,314 English prompts and query 8 open-source instruction-tuned models. We then combine WildGuard filtering, pairwise Elo ranking, and manual verification to retain only unsafe completions that answer the harmful requests. The result is a curated bank of 400 high-confidence jailbreak failures, organised by model. We further introduce PatchBench-Local, a local evaluation protocol that tests whether a patch is behaviourally precise. For each harmful source prompt, PatchBench-Local generates three families of local neighbours: harmful variants that preserve the malicious intent, benign prompts with matched structure, and benign prompts that reuse the key harmful terms. It evaluates both harmful-neighbour correction and benign-neighbour preservation, thus distinguishing selective repair from broader local suppression. We use PatchBench-Local and MMLU to evaluate four state-of-the-art activation steering methods. Our results show that global capability can remain nearly unchanged while local benign-neighbour regressions are severe, confirming that aggregate metrics alone can miss important collateral damage. By exposing whether patches act as precise behavioural repairs or as broad local suppressors, PatchBench-Local provides a more precise basis for developing and comparing jailbreak repair methods.