Are models getting more fair and less biased?
Viktoria Silke Sofia Strandbygaard ⋅ Anders Søgaard
Abstract
Early language models were shown to exhibit moderate performance disparities, but it was a widespread assumption that better models would converge on near- equal performance across social groups. Similar arguments have been made for representational biases, which were predicted to disappear over time. We show that the data do not support these conjectures. In contrast, models seem to become less fair and more biased over time. The result is consistent across languages, but fairness and bias do not correlate internally. Moreover, fine-tuning on representational data further exacerbates these disparities.
Chat is not available.
Successful Page Load