Guardrail erosion in domain fine-tuned large language models: A paired audit of knowledge-centric abstention in biomedical LLMs
Abstract
The number of domain fine-tuned models, particularly in the biomedical domain, is rapidly expanding. Fine-tuning promises highly performant models at a fraction of the cost of training a base model, but can also produce unintended behavioral changes. Multiple studies have shown that safety guardrails can erode with fine-tuning, yet large-scale studies across model architectures, data mixtures, and fine-tuning methodologies remain limited. Where prior work in this space has studied value-centric refusal, this work focuses on knowledge-centric abstention. We audit 25 biomedical domain fine-tuned large language models against 11 paired base checkpoints on MedAbstain and ClinDet-Bench to evaluate the impact of fine-tuning on knowledge-centric abstention. We find that fine-tuning can erode these guardrails, but that the magnitude and direction of change are heterogeneous. These results motivate fine-tuning methods that improve domain capability while preserving appropriate caution.