ML-assisted Randomization Tests for A/B Experiments
Wenxuan Guo ⋅ JungHo Lee ⋅ Panos Toulis
Abstract
We study the problem of detecting treatment effects in randomized A/B experiments when the effects are potentially small and complex. A common approach in such settings is to fit machine learning (ML) models of outcomes on observed features and then apply classical $t$-tests to residualized outcomes. We show that this approach can suffer substantial power loss under treatment effect heterogeneity. We propose a randomization-based test whose statistic measures the out-of-sample predictive gain from including the treatment variable in a flexible ML model. Leveraging experimental randomization and sample splitting, our test is finite-sample valid for arbitrary models and loss functions. We establish power guarantees under heterogeneous treatment effects and demonstrate substantial empirical gains in simulations and large-scale A/B experiments.
Chat is not available.
Successful Page Load