Break It to Align It: Annotation-Free Medical Safety Alignment via Preference Inversion
Abstract
Safety alignment in specialized domains such as medicine requires preference data that pairs a safe response with a plausible unsafe one on the same prompt. For ordinary clinical questions no such data exists, and constructing it conventionally requires domain experts to author or validate harmful medical advice. We propose Adversarial Self-Correction (ASC), which obtains these pairs from a general-domain safety preference dataset alone. ASC first trains two policies from a single backbone on the same data, differing only in the order of each preference pair. Inversion raises lexical toxicity only from 0.099 to 0.149 while raising compliance with unsafe requests from 10.5\% to 83.0\%, so what it recovers is unsafe compliance rather than profanity. ASC then samples one response from each policy for the same unannotated medical question under an identical generation context, orders the pair by provenance, and runs DPO from the safety-aligned policy; no teacher, judge, or domain annotator is used during pair construction. On BioMistral-7B, ASC reduces unsafe compliance from 10.5\% to 5.0\% beyond general safety training and lowers MedSafetyBench harm from 1.31 to 1.23, at a cost in safe-prompt compliance. On the already aligned Llama-3.1-8B-Instruct the effect is close to neutral, consistent with the gain depending on how much unsafe behavior remains to remove. Opposite orderings of a single preference dataset thus yield a generator of domain-specific safety preference pairs without any domain-specific annotation.