Less Abstention, More Bias? Demographic Bias Drift After Supervised Fine-Tuning
Seema Guruvadoo ⋅ Pavithra Nair ⋅ Gilad Gressel ⋅ Krishnashree Achuthan
Abstract
Supervised fine-tuning, even on seemingly benign data such as math or code, is known to increase bias in large language models (LLMs). Yet knowing that bias increases tells us little about what fine-tuning is actually changing: which biases are affected, whether the same biases emerge across models and domains, or what behavioral changes accompany them. Understanding this structure is critical for detecting unintended effects of fine-tuning and, importantly, for determining how they can be prevented. We systematically map demographic bias before and after fine-tuning 3 LLMs on 12 datasets spanning 4 domains: math, code, medicine, and safety. Across 36 fine-tuning experiments (12 datasets $\times$ 3 models), demographic bias increased substantially in 28 cases, with the mean post-fine-tuning bias score reaching $2.7\times$ its baseline value. The increase in bias follows a consistent demographic pattern: socioeconomic status, physical appearance, and sexual orientation show the largest relative increases (2.4–4.8$\times$ the baseline), while bias in race/ethnicity and religion categories remain essentially unchanged. Underlying this pattern is a consistent reduction in abstention (choosing not to answer when the evidence is insufficient): after fine-tuning, models are more likely to answer ambiguous demographic questions even when the available information does not support a definite answer. To test whether preserving abstention during fine-tuning can prevent the increase in demographic bias, we repeat fine-tuning while replacing $\sim$10\% of the training tokens with items from the CoCoNot dataset, for which the appropriate response is to abstain or ask for clarification. Among the 10 experiments where standard fine-tuning increased bias, adding abstention data to the fine-tuning dataset reduces this increase in 9, while nearly eliminating it for socioeconomic status and physical appearance and reversing it for age. These results highlight the need for evaluating fine-tuned models for unintended changes beyond target-task performance and identify abstention as an important property to preserve during fine-tuning.
Chat is not available.
Successful Page Load