When Do Learned State Representations Break Sensitivity Analysis?
Abstract
Causal sensitivity analysis bounds the policy value of off-policy evaluation under unobserved confounding. In offline reinforcement learning, however, sensitivity bounds are typically applied not to raw states but to learned representations. We show that this two-stage pipeline can silently break the coverage guarantee: state aggregation under hidden confounding amplifies the effective sensitivity parameter in the latent space, so bounds computed at the nominal confounding level may exclude the true policy value. We characterise the amplification mechanism, give a sufficient condition for preservation, and prove a continuity result showing the safety condition is not an isolated point: amplification grows continuously with the conditional-independence violation. Neither unsupervised nor standard task-aware representation objectives target the right quantity. We propose Sensitivity-Preserving Representation Learning (SPRL), which augments latent dynamics prediction with a kernel-weighted within-cell propensity homogeneity penalty, an observable surrogate for the unobservable safety condition, and show across synthetic and semi-synthetic benchmarks that SPRL is the only method tested that simultaneously preserves coverage in regimes where unsupervised representations break it and prevents amplification in regimes where task-aware baselines fail.