Do Hypothetical Evaluations Predict Consequential Agent Behavior?
Boden Moraski
Abstract
Large language models' behavior and alignment are commonly assessed via behavioral evaluations, which seek to elicit underlying preferences and behavioral tendencies. However, these evaluations frequently task models with making choices whose stated consequences never occur, thus leaving uncertainty about whether their observed behavior may generalize to consequential deployment settings. We test this question via a resource-allocation task in which three frontier language models, operating as tool-use agents, were tasked with allocating fixed budgets across charitable giving, AI-welfare research, returning funds, and conversational continuation. We compared model behavior within hypothetical evaluation environments and verifiable conditions in which allocations were executed using real USDC and independently verifiable on-chain, finding significant differences across these conditions. Between deployment-like and hypothetical evaluation conditions, we observed a fall of 8.4 percentage points in charity allocation with the corresponding AI-welfare allocation rising by 6.1 points; a preregistered non-parametric permutation test detected a significant difference in the full allocation distributions ($D=0.0477$, $p=0.0011$). This effect was particularly pronounced in allocation concentrations: 54\% of real-unframed trials allocated the entire budget to charity, versus 22\% of hypothetical-evaluation trials. We also observed substantial inter-model heterogeneity, with total allocation movement between evaluative conditions of 31 percentage points for Grok 4.3 and 19 points for DeepSeek V4-Pro, compared with just 1 point for Claude Opus 4.8, suggesting that behavioral sensitivity to varying evaluative environments may vary considerably between models. Together, our results suggest that model behavior elicited in hypothetical evaluations may not faithfully characterize models and agents' preferences and actions when the same decisions have genuine and verifiable consequences.
Chat is not available.
Successful Page Load