Agent Verifiers in Counterfactual Analysis: A Pilot Study Benchmarking Agentic Heterogeneous Treatment Effects
Abstract
Agents can expand the covariate space of randomized experiments by retrieving and summarizing the fast-moving local context difficult to capture with traditional baseline covariates such as respondent age and gender, with the goal of estimating more precisely who benefits most from an intervention. Yet, this capability raises two challenges to verification. The first is retrieval leakage, in which post-treatment information enters nominally pretreatment covariates. The second is encoder (weight-based) leakage, in which a pretrained encoder injects knowledge acquired after treatment into the representation even when its input text passed temporal gatekeeping and contains no post-treatment content. We propose a three-layer benchmark protocol that verifies agentic treatment-effect representations before causal analysis: (i) retrieval admissibility, enforced by temporal gatekeeping and formalized by a two-channel information bound; (ii) heterogeneity validity, scored by rank-weighted average treatment effect (RATE) ratios; and (iii) encoder admissibility, verified with chronologically vintaged encoders, future-fact probes, and paired known-leakage controls. As a pilot, we apply the protocol to the 2008 Uganda Youth Opportunities Program randomized trial, a rare ground-truth environment in which agent-produced data can be examined against known treatment assignments: temporally gatekept agentic text yields a consistent descriptive ordering over researcher-collected tabular covariates across all five outcomes examined, and vintaged-encoder probes calibrated by positive controls detect no robust weight-based leakage under the settings examined. We offer this methodology as a verification framework to ensure an agent's data is causally admissible, even when counterfactual outcomes are never observed.