Auditing Process-Level Validity in LLM-Based Social Simulation: A Case Study of the FOMC
Abstract
LLM-based social simulations are usually evaluated by whether their aggregate outputs agree with observed outcomes. We audit a persona-grounded simulation of the U.S. Federal Open Market Committee (FOMC) at four levels (outcome agreement, member composition, artifact fidelity, and procedural coherence) and report run-to-run variability at each, using three redraws of a 17-meeting retrospective window, 30 corpus-audit runs, and 18 simulations of the July 2026 meeting completed under a July 15 corpus cutoff before the decision was released. The levels come apart. In the July set, 13 of 18 draws record the realized hold, or 11 if direction is read from the generated Statement. The member-level results are negative: both architecture configurations score below an always-assent baseline on vote alignment (82.8% and 77.7% versus 93.6%), and dissenter-set overlap is indistinguishable from chance (Jaccard 0.18 versus 0.16, p = 0.18). The most persistent simulated dissenter favors a hike in every hold-motion draw, contrary to her vote at the July 2026 meeting, and the pattern is already present in the pre-deliberation beliefs. Three LLM judges score Statements at 3.86/5 and Minutes at 3.28/5, with fact alignment lowest for every judge; a crossed generator–judge probe finds no same-provider preference, while judge-free checks find 11 of 18 July Statements naming a former Chair. Procedural checks find defects in 8 of 13 rate-moving runs, each in the emitted record. Agreement at one level is therefore weak evidence about the others, and repeated runs make the variation within each level visible.