The Survey Item Predicts Human Disagreement Better Than Simulator Output Entropy
Anqi Peter Li ⋅ Kundana C Kommini ⋅ Ethan Yip
Abstract
Social simulators are increasingly used to model populations, but varied generations do not show that a simulator knows which questions divide people. We study this prediction problem on GlobalOpinionQA, where the target is the dispersion of human responses. A frozen $335$M-parameter sentence encoder that reads only the rendered survey item and never runs the simulator predicts this dispersion better than the corrected output entropy of a $7$B-parameter simulator. On held-out questions, the encoder reaches Pearson-$r=0.393$, versus $0.200$ for output entropy: a margin of $+0.194$ $[+0.146,+0.240]$. The ordering holds in all five instruction-tuned model families we evaluate. We then compare information in the rendered item, hidden states, and the emitted option distribution. A latent readout recovers dispersion, and a full-distribution readout improves total-variation distance at every option count; temperature alone cannot alter option ranking. Within this controlled forced-choice interface, the evidence supports an item-level text-versus-output comparison rather than a claim about full social simulation.
Chat is not available.
Successful Page Load