Mechanism Oriented AI based Simulation for Predicting Policy Effects in Peer Review Conferences
Abstract
AI agents powered by large language models enable simulations of peer-review conferences through interactions among reviewers, authors, and decision makers. Such simulations could help organizers anticipate the consequences of policy changes that are difficult to test in real conferences. However, matching historical ratings or generating plausible reviews does not establish that a simulator will reproduce human responses to new rules. We formulate conference simulation as a mechanism-oriented problem, extending pattern-oriented modeling to policy responses across multiple levels of review behavior. The framework requires explicit channels through which policies affect outcomes and evaluation metrics that capture both observed changes and demonstrated stability. Guided by these principles, we develop a multi-agent simulator that combines language-based interactions with data-driven calibration. A verifier converts interactions into typed evidence, and calibrated transitions fitted on years disjoint from evaluation translate that evidence into rating updates. We evaluate the simulator against human responses to the ICLR 2019--2020 changes in rating scale and discussion duration. Our system matches the direction of all six evaluated responses and the magnitude of four of the five responses with estimable uncertainty, outperforming existing systems. It is also the only evaluated system that avoids inventing an effect on the metric where humans remained stable; a policy-as-prompt baseline over-predicts that effect by a factor of 27.These results support mechanism-oriented modeling and policy-response validation as a basis for assessing conference simulations intended to inform policy decisions.