Who Keeps Giving Feedback? Participation Feedback Creates Alignment Lock-In
Abstract
A deployed alignment policy can change who supplies its next round of feedback. If better service elicits greater participation from one group, feedback-based updates can amplify an initial imbalance even when preferences do not change. Building on prior models of performance-dependent user retention, we study representative recruitment as an intervention in this loop. A two-group model yields an exact threshold: the representative policy is globally attracting when the feedback gain is at most 1, but two unrepresentative attractors appear above that threshold. At the equality boundary, convergence remains stable but becomes polynomially slow. We derive how unequal population shares, biased audit feedback, delayed responses, and finite feedback change this conclusion. Simulations separate across-deployment averages from within-deployment lock-in and show when a local noise approximation is informative. The resulting measurement protocol distinguishes a policy's effect on response propensity from its effect on preferences. Representative feedback is not just an evaluation sample: its effective weight and sampling quality can determine the stability of the alignment process itself.