Calibrating What Behavior Can Identify: Conformal Reward Sets for Inverse Reinforcement Learning
Yuhan Chi
Abstract
Inverse reinforcement learning (IRL) returns a reward, and a policy is then optimized against it as if it were correct. Conformal inverse optimization (CIO) offers a principled alternative for one-shot inverse problems: calibrate a radius around the point estimate on held-out decisions, then optimize robustly over the resulting set. We ask what survives the move to IRL. The conformal step transfers untouched; two others break, both because behavior pins down a reward only up to transformations that leave every policy value unchanged. First, CIO scores a demonstration by an angle between reward vectors, which depends on coordinates the problem never fixes: a feature reparameterization that changes no reward and no demonstrator's policy still moves the calibrated angular radius by $24^\circ$ on average. Scoring instead by the spread a reward difference induces over achievable policy values restores invariance and keeps the score a linear program. Second, which event a quantile certifies is decided by which distance to the cone of rewards explaining new behavior is scored: the nearest point gives intersection with that cone, the farthest point of the normalized cone gives containment of the latent reward. Only the first is a linear program, and at a $0.80$ target we measure the two events at $0.82$ and $0.44$. We give the enlarged radius that prices the relaxation, and the point past which the resulting robust program provably says nothing.
Chat is not available.
Successful Page Load