Fine-Tuning Fixes Mode Collapse and Over-Dispersion in LLMs
Abstract
Recent work raises concerns that LLM outputs tend to be under-diverse. Our results show that this phenomenon is model- and dataset-specific, and that with sufficient fine-tuning data, LLM output diversity converges toward the diversity of the training data. We say that a model is mode-collapsed when two independent samples are more likely to coincide exactly (collide), or to be similar under a kernel, than two independent draws from the target distribution---often a human population---that the model is trained to represent. We derive a bias--variance decomposition of the expected gap between the model's and target's collision probabilities, showing that finite-sample SFT can leave a model either under- or over-dispersed. Finally, we show that the absolute gap between model and target collision probabilities is bounded by the square root of the excess cross-entropy minimized by log-loss training. Consequently, a model sufficiently close to optimal under log loss cannot exhibit arbitrarily miscalibrated diversity. We test the decomposition and the bound in three experiments: (1) we train 100 GPT models across 100 log-spaced sample sizes on two synthetic 16-gram languages; (2) we LoRA fine-tune four LLMs on responses from three sociological surveys (GSS, ANES, WVS); and (3) we repeat (2) on CodeNet, a dataset of human solutions to coding tasks, measuring diversity with a similarity kernel over program abstract syntax trees. We find substantial heterogeneity in over- and under-diversity across models and datasets. Fine-tuning moves relative diversity toward the human (or synthetic target) level in all experiments, consistent with our theoretical predictions. Diversity miscalibration can arise from finite-sample error and shrink as supervised fine-tuning better approximates the target distribution. Across our survey and code experiments, increasing target-distribution data moves model diversity toward the target level. We call for future research into post-training methods that make use of the diversity-calibrating properties of fine-tuning.