From Observation to Intervention: Identifiability of Latent User Bias in Human–AI Interaction
Lucas Biechy ⋅ Yuxiao Li ⋅ Zhonghao He ⋅ Tianyi (Alex) Qiu
Abstract
Alignment methods typically treat user behavior as a reliable proxy for underlying preferences. However, this assumption breaks down when users hold belief-dependent cognitive biases, risking irreversible ``confirmation traps'' and harmful feedback loops. In this paper, we formalize when an AI system can distinguish true user preferences from behavior shaped by cognitive bias. Modeling this interaction as a multi-armed bandit, we first prove a non-identifiability theorem: under passive observation, a static binary biased user produces action trajectory distributions strictly identical to an unbiased user in a transformed environment, forcing any action-only estimator to perform no better than random guessing. We then show that active preference elicitation breaks this observational equivalence. Specifically, we prove that an Observer applying a finite sequence of adversarial zero-reward interventions of length $m_\delta$ makes the post-intervention action distributions arbitrarily separable.Our theoretical results suggest that robust alignment requires actively intervening to tell informed preferences from biases.
Chat is not available.
Successful Page Load