Strategic Causal Policy Learning: Welfare, Safety, and Fairness Thresholds
Abstract
The standard rule in unconstrained causal policy learning is simple: treat individuals with nonnegative CATE. We show that this rule can fail when treatment eligibility is strategically manipulable. Agents may move or report covariates after a policy is announced, so treatment is assigned using manipulated covariates while welfare remains determined by baseline CATEs. This creates a mismatch between who receives treatment and who benefits from treatment. We formalize this setting as strategic causal policy learning and show that strategic welfare equals the CATE-weighted mass of baseline types that can reach the treatment region. This access-burden representation leads to three distinct strategic thresholds: a welfare threshold, a safe threshold, and a fair threshold. For linear CATE models, we derive closed-form threshold corrections. Homogeneous costs admit a single buffer correction that recovers the ideal positive-CATE allocation, while heterogeneous costs make welfare maximization, no-exploit safety, and subgroup fairness select different policies. Experiments illustrate these tradeoffs and show when group-specific corrections can restore the ideal allocation.