Two Directions of Adaptation: Measuring Human-to-AI and AI-to-Human Influence in Coupled Systems
Abstract
Repeated human--AI interaction changes both participants. Standard alignment evaluation scores whether the model follows a fixed human preference. Agreement can rise while the person's independent choice set contracts, confidence becomes dependent on the model, or the model learns a shallow way to obtain approval. We propose a paired measurement for coupled systems. One curve records how much bounded human intervention moves the model's state; the other records how much bounded AI intervention moves the human's state. Matched trajectory replays estimate the cost-normalized effect in each direction, a retention test after the AI is withdrawn separates learning from dependence, randomized advice validity separates evidence-based updating from overreliance, and induced preference shift is measured against private values recorded before exposure. Executed studies already show that AI help raises performance while present and lowers it after withdrawal, and that opinionated assistants move users' views. The proposed design measures those effects as functions of intervention budget, in both directions, with preregistered estimands.