Before the Feedback Loop: Preference-Span Shift in Dynamic Human–Artificial Intelligence Alignment
Abstract
Dynamic alignment assumes that a system can interpret new feedback before deciding how to update. We isolate a failure that precedes the update rule: a low-rank personalized reward model can represent a new user only within preference directions learned from its training population. Random held-out-user splits can therefore certify identity transfer while missing preference-support shift. In a controlled benchmark, LoRe (low-rank reward modeling) falls from 75.7% in-span accuracy to 51.1% at full shift; a 12-parameter residual reaches 68.3%, while a 16-parameter correction reaches 73.5%. Public PersonalLLM reward vectors reproduce this controlled capacity pattern. Offline labels from the participant-indexed PRISM Alignment Dataset reveal an opposing statistical constraint: with four support comparisons and 1,024-dimensional reward-model features, the tested full user correction does not improve global Bradley-Terry. An exploratory 12-parameter, globally anchored and validation-shrunk residual improves cluster-holdout accuracy from 61.8% to 65.1% in 9 of 10 pipeline runs, with no higher mean negative log-likelihood; paired tests disagree. Always-on adaptation reaches higher accuracy but severely worsens negative log-likelihood. These results motivate representational readiness as a prerequisite for coupled alignment: evaluate preference-support shift, compare estimable capacities, preserve a population fallback, and use confidence-sensitive scoring. This offline primitive does not model reciprocal human change, but it can reduce the risk that a feedback loop mistakes representational blindness or few-shot overconfidence for progress.