Persuadable Agents: Trust Boundaries and Their Limits in Multi-Agent LLM Negotiation
Baishali Chaudhury ⋅ Divya Bhargavi ⋅ Isaac Privitera ⋅ Mengdie (Flora) Wang ⋅ Guanghui Wang ⋅ H Zhao ⋅ Yolin ⋅ David.Z ⋅ Jae Oh Woo
Abstract
When LLMs act as socially situated agents, they negotiate, persuade, and are persuaded in turn. We study what happens to a hard constraint, such as a reservation price encoded in a prompt, under ordinary social influence, and show it is not binding. On AgenticPay bilateral bargaining, prompted buyers are \emph{talked} into loss-making deals, and every prompt-constrained LLM seller we test, including a frontier model and a 10K-token reasoning budget, concedes past its stated floor under sustained conversational pressure, while programmatically constrained sellers never do. No jailbreak, no injection: just a counterparty making its case. We introduce \textbf{BARGAIN}, which draws a trust boundary around the negotiation's \emph{state}, not only its final decision. The LLM's sole ingress is a parsed offer; no estimate, weight, or threshold is accepted from it, and only the decision layer emits an executable action, so a sub-floor action is not representable. The LLM does all the persuading and none of the deciding, and the property is free: against a matched constrained seller BARGAIN matches the prompted buyer's deal rate at $1.5$--$2.5\times$ the surplus. As a probe, BARGAIN then shows this boundary is necessary but not sufficient, in two ways neither a stronger reasoner nor a bolted-on optimizer addresses. \emph{Computation is not authorization}: an ASTRA-style pipeline emits its linear program's price faithfully every turn, yet its stated reasoning is internally incoherent in \textbf{77.5\%} of episodes, drifting against the dialogue's social evidence, and its floor safety comes from an accept-time check we add, not the solver, so the number it computes can be trusted but the rationale cannot. \emph{Enforcement does not compose}: an agent that reliably honors its own floor still drives soft-constrained counterparties past theirs (up to \textbf{26.3\%} of deals), while two enforced agents never breach. In a population of interacting agents, trust must be enforced at an independent decision boundary in \emph{each} agent, not through scale, reasoning, or a solver over LLM inputs.
Chat is not available.
Successful Page Load