Flexible but Fragile: Budgeted Human--Digital-Twin Allocation for A/B Tests
Abstract
LLM-based digital twins (DTs) use an individual’s survey or interview history to predict how that person would respond under a new experimental condition. These predictions can provide low-cost synthetic outcomes, but they may be systematically biased, and that bias must be estimated from a small pilot sample in which both human and DT outcomes are observed for the same individuals. We study how to allocate a fixed experimental budget between human and DT outcomes to estimate the ATE of a broader target population from a limited accessible panel. We compare three allocation policies that differ in how much flexibility they allow and how they account for pilot uncertainty. A constrained policy fixes the human sample before selecting DT outcomes. A joint policy selects human and DT counts together. An uncertainty-aware policy retains this flexibility while penalizing uncertainty in the estimated ATE bias. We evaluate the three policies on log willingness to pay in the Digital Certification substudy of the Twin-2K-500 follow-up mega-study, across 60 combinations of pilot size, DT cost, and deployment budget, using 100 pilot–holdout splits per configuration. Greater allocation flexibility alone does not improve out-of-sample performance. The constrained policy has significantly lower holdout MSE than the joint policy in 51 of 60 configurations, with no significant difference in the remaining nine. Penalizing bias-estimation uncertainty reverses this pattern in the main evaluation. The uncertainty-aware policy significantly outperforms the joint policy in 46 configurations and the constrained policy in 40, while underperforming the constrained policy in six.