Path-Specific Counterfactual Preference Alignment: A Causal Framework for Dynamic Human–AI Systems
Abstract
Alignment methods that learn from human feedback typically treat preferences as an exogenous target. In long-horizon human--AI interaction, however, the system's actions can alter the preferences that later generate its feedback. Penalizing all AI-induced preference change is too restrictive: information, education, and reflective deliberation can legitimately change what a person wants. We introduce \emph{Path-Specific Counterfactual Preference Alignment} (PS-CPA), a causal framework that evaluates preference trajectories relative to a reference interaction while preserving designated admissible pathways and neutralizing designated inadmissible pathways. We formalize dynamic preferences in a structural causal model, define a reference-relative path-specific preference trajectory, and distinguish total AI influence from influence transmitted through normatively restricted mechanisms. We show that a path-specific objective removes direct instrumental incentives to steer preferences through blocked pathways under explicit graphical conditions, derive a stability bound for accumulated influence in linear dynamics, and bound objective error under imperfect counterfactual estimation. We then give identification conditions, discuss failure under unmeasured confounding and recanting-witness structures, and outline estimators based on longitudinal g-computation and latent state-space models. Finally, we propose synthetic and matched-content human-interaction evaluations that separate informational learning from manipulative framing. PS-CPA does not infer which influence is morally permissible; rather, it supplies causal machinery for enforcing an explicitly specified partition of admissible and inadmissible influence in co-evolving human--AI systems.