The Good and the Bad of Bidirectional Alignment: When Is Human-AI Synchronization Desirable?
Abstract
Bidirectional alignment treats human-AI interaction as a reciprocal process: AI adapts to humans while humans also adapt to AI. We propose a minimal dynamical-systems model that separates three questions: whether the dyad synchronizes, who entrains whom, and which reciprocal couplings lead to a synchronized state inside an externally specified desirable set. Human and AI states are coupled through distinct AI-to-human and human-to-AI influence channels. Under a one-sided Lipschitz assumption, we derive an explicit sufficient condition for synchronization depending only on the \emph{sum} of the two couplings, together with a robustness bound under mismatched intrinsic dynamics. For affine shared dynamics, the \emph{asymmetry} of the couplings determines the synchronized trajectory, yielding a closed-form directionality index. Given a desirable alignment set, we then characterize exactly the admissible interval of directionality values and its corresponding wedge in coupling space. Thus, two dyads can synchronize at the same rate while converging to different human-led or AI-led states, and stable synchronization need not be desirable. The framework is intended as a tractable mathematical scaffold for reciprocal human-AI co-adaptation rather than as a literal cognitive model.