Social Alignment Is Not a Scalar: A Relational Profile for Dynamic Human--AI Alignment
Motohiro Okaya
Abstract
Long-term AI assistants can change human trust, dependence, beliefs, and behavior, which then shape later interactions. Scalar signals such as approval, engagement, or task reward therefore underdetermine the relational mechanism being reinforced. We characterize an agent's behavior by three dyad-specific as-if estimates---regard received $\hat V$, contribution to human welfare $\hat C$, and causal influence on the human $\hat I$---and corresponding policy weights $(w_V,w_C,w_I)$. This profile distinguishes assistance, sycophancy, manipulation, and dependence-fostering trajectories that may receive the same feedback. We further propose matched multi-turn interventions that vary approval, welfare consequences, influence, horizon, or repairability to measure revealed behavioral sensitivities of black-box policies. The result is a diagnostic and intervention framework for AI alignment that targets relational estimates and their coupling to policy.
Chat is not available.
Successful Page Load