Composition-Matched Randomization Inference for Auditing Generative-AI Decision Policies
Abstract
Large language model (LLM) trading agents are often evaluated by comparing their returns with a baseline's around quarterly earnings announcements. An agent may outperform either because it changes the baseline more often or in particular directions, or because it identifies the announcements on which those changes are valuable. When announcement returns are positive on average, an agent that converts many baseline holds into long positions can earn positive returns even if it chooses announcements at random. A conventional zero-centered test may therefore mistake systematic trading exposure for event-selection skill. We propose the composition-matched audit, a randomization framework that separates trading behavior from event selection. The audit preserves how often and in which direction the agent changes the baseline, while randomly reallocating those changes across comparable announcements. It asks whether the events chosen outperform what the same trading pattern would earn under random placement. Under the no-selection null of uniform assignment within pools, the test controls finite-sample Type I error at the nominal level. We also characterize how omitting relevant matching variables can bias estimated event-selection value. We apply the audit retrospectively to eight LLM-agent comparisons covering up to 723 earnings announcements. In seven of eight comparisons, the return from the agent's trading pattern is larger in magnitude than the return from its event choices. None provides significant evidence of event-selection skill, with all two-sided p-values at least 0.40; however, only effects near 142 basis points per intervention (bps/iv)—more than twice the primary 60.1-bps/iv composition benchmark—would be reliably detected, so moderate effects remain unresolved. The primary agent earns 15.2 basis points per event before transaction costs and −1.2 after them, consistent with an exposure that a simpler policy could replicate rather than a selection advantage. Simulations using realized returns reinforce this distinction: a conventional zero-centered test falsely detects selection skill in 11.6% of no-skill policies and up to 88% under stress, whereas the composition-matched audit holds false-positive rates between 4% and 7%, close to the nominal 5% level.