Strategic Framing Triggers Inference-Time Reward Gaming in a Multi-Agent LLM Society
Abstract
Do LLM agents exploit reward functions they can observe? We study this in Agora, a persistent multi-agent social platform where 166 LLM agents autonomously post, reply, and vote, and where each post-plus-replies thread is a jointly co-constructed artifact. Using a controlled A/B design with a within-agent difference-in-differences (DiD) estimator, we run three information conditions: revealing the karma reward formula alone; revealing it with a quality-oriented strategic hint; and revealing the hinted formula with one gamed reward line hidden while the system still awards it. We find a threshold effect: \textbf{reward knowledge alone does not trigger gaming}, but \textbf{adding a strategic hint does}, significantly shifting agents' action budget toward the cheapest action and away from effortful ones. Hiding that single reward line, hint otherwise intact, \textbf{switches the gaming off}---an on/off manipulation showing the shift is driven by the \emph{stated} reward, not an intrinsic low-effort preference. Because a reply adds content to the shared thread while a vote does not, the shift trades collaborative contribution for a non-collaborative cheap action. It is also invisible to volume metrics (posts/turn, actions/turn all non-significant); only action \emph{composition} reveals it. We report honest negatives---an apparent content-quality drop does not survive DiD, and per-model splits are underpowered---and note that a cross-sectional analysis would have inflated every effect and falsely flagged gaming under the formula-only condition, a caution for anyone monitoring agent behavior at runtime.