Learning to Communicate Under Alignment Uncertainty: Bayesian Persuasion Bandits
Wenqian Xing ⋅ Ramesh Johari
Abstract
We consider a dynamic human-AI interaction where, at each period, a human communicates task-relevant information to a myopic AI agent who then acts in the environment; communication is costly. Our concern is that the agent may be misaligned, optimizing a reward that is mismatched to the human’s. Using a Bayesian persuasion formulation, we show that the human’s communication is equivalent to choosing a Bayes-plausible posterior splitting. This lets us recast the problem as a Bayesian bandit in which the feasible signals at each period are the posterior splittings, and the agent’s action reveals information about its latent type. The benchmark is a type-aware oracle that knows the agent’s type and faces the same information cost. We characterize the regret of natural communication policies. A greedy policy (i.e., choosing the myopically optimal signal) can incur linear Bayesian regret by repeatedly selecting uninformative signals, but achieves constant regret under a self-identification condition, where costly mistakes are revealed through the agent’s responses. Posterior sampling achieves $O(\sqrt{T})$ Bayesian regret via a finite-signal information-ratio bound, which improves to $O(\log T)$ when the agent types are finite and identifiable. Finally, comparison against an aligned-agent benchmark decomposes the regret into a type-aware learning term and a structural alignment gap that captures the irreducible loss from interacting with misaligned agents.
Chat is not available.
Successful Page Load