Toward High-Fidelity LLM Simulation: Lessons from a Mobilization Experiment
Abstract
LLM-based agent simulations offer a new way to study interventions in complex systems, but their role in decision-making and policy evaluation remains unclear. We examine this question using LLM-SocioPol, a social-media environment with 20,000 heterogeneous LLM agents, benchmarked against a large voter-mobilization field experiment. We compare an informational message that provides a general appeal to vote with a social message that adds aggregate and peer-specific cues about other users’ voting intentions. The simulator reproduces the qualitative ordering observed in the field experiment, with a stronger response to the social message, but generates substantially larger effect sizes. These results illustrate both the flexibility and the limitations of LLM-based simulation. The environment permits repeated counterfactual comparisons and controlled study of experimental designs, but a fixed modeling recipe does not establish behavioral validity arising from population construction, interaction rules, and prompts. Rather than offering a definitive prescription for how such simulators should be constructed or used, our results point to a broader methodological opportunity: developing useful simulators requires starting with the intended inferential or decision goal and then aligning the population data, interaction structure, behavioral mechanisms, and validation strategy with that goal. This perspective suggests looking beyond prompt tuning alone toward the joint design and evaluation of the full simulation system.