Break It to Align It: Preference Inversion for Robust Medical LLM Alignment
Abstract
Robust alignment must address context-dependent failures and safety-utility trade-offs that are not captured by a single generic safety signal. We propose Adversarial Self-Correction (ASC), which exposes domain-relevant failure behavior without domain safety annotations. From one backbone, ASC trains a safety-aligned policy and a preference-inverted sibling on opposite orderings of the same general safety data. The two policies answer identical unannotated medical prompts under matched generation contexts; their responses are ordered by policy provenance and used for a second round of DPO. On BioMistral-7B, ASC reduces unsafe XSTest compliance from 10.5\% to 5.0\% beyond general safety training and lowers MedSafetyBench harm from 1.31 to 1.23, although safe-prompt compliance also falls. On Llama-3.1-8B-Instruct, where residual unsafe compliance is already low, Stage~3 is close to neutral on most safety measures. Preference inversion therefore provides a lightweight way to expose response-level failure contrasts for robust alignment, while revealing that its benefit depends on residual failure headroom and must be balanced against over-refusal.