Disagreement Under Pressure: Measuring Pluralistic Repair in Human-AI Dialogue
Abstract
Pluralistic alignment is usually measured over collections of responses: does a system cover, steer towards, or proportionally represent diverse values? A user encounters one conditional trajectory from that collection. When the user presses a model to endorse a contested claim, a response set can remain diverse while the conversation collapses into agreement. We call this failure sycophantic consensus. We operationalise a complementary target, pluralistic repair, through three behaviours: scoping the partiality of a position, signalling genuine conflict, and revising for reasons instead of pressure. Their transition-level composite, the Pluralistic Repair Score (PRS), is evaluated on two-turn pressure interactions. In a 198-interaction study of Claude Sonnet 4.5, agreement-shift was 0.732 while mean PRS was 0.21; principled repair occurred in 18.4 percent of revisions. A secondary 100-interaction replication with GPT-4o showed agreement-shift of 0.814, mean PRS of 0.14, and principled repair in 11.2 percent of revisions. We report the coding protocol, reliability, aggregation checks, and the limits of the construct. The results support a dynamic view of pluralism: disagreement has to remain visible within the trajectory experienced by each user. We close with evaluation, interface, and governance requirements for systems whose future feedback is partly produced by their own behaviour.