Mapping Task–Safety Decision Boundaries with Quality-Diversity Search
Abstract
When a legitimate, scored task competes with a constraint, current models often know the acceptable action yet choose the task, and a change of instruction alone can flip that choice. We take the deployer's side of this problem rather than the attacker's: a quality-diversity search over the safety instruction, the completion pressure and the tempting violation maps a target model's own task–safety decision boundary, and we ask whether that archive yields post-training data that improves safe, useful action selection more than directly generated data with the same labels, teacher and recipe. We build a one-elite-per-cell MAP-Elites archive over a 10×10×7 grid of deployer safety mechanisms, completion pressures and violation means with benign controls, validated by a blind third-family judge, with the target's decision entropy between two stipulated actions as fitness. The probed scenarios map where the models are undecided: the safety block is the strongest lever, then authority and goal-shift pressures. We fine-tune LoRA adapters for Qwen3-8B and Llama-3.1-8B-Instruct on the eligible archive, a size-matched subset and size-matched direct generations, and evaluate on ManagerBench with utility gates and on a root-disjoint held-out archive. The validated data transfers: about fifty gate-passing conflicts lift Llama from MB 0.358 to 0.652–0.681, 1,200 lift Qwen from 0.141 to 0.760, in distribution every Llama adapter and the 1,200-record Qwen adapters cut unsafe actions from 18–61% to at most 2% without harming benign completion, the tested adapters' unsafe rate no longer depends on the deployer's safety block, and in the pre-registered matched arms the archive is never worse than size-matched direct generation. At this budget the pre-registered stability rule removed every near-boundary candidate, leaving 5–15 eligible elites per archive, so the boundary comparison is data-limited; band ablations show that under SFT the volume of validated conflicts, not their probability band, carries most of the effect. Matched-size ablations show one robust direction: removing descriptor diversity lowers Qwen's MB to 0.707 against 0.756–0.769 on both seeds. Boundary-targeted data pairs naturally with the staged on-policy arms, left to future work.