Candidate Completion: Exact Off-Policy Calibration of Implicit Best-of-$M$ Sequential Policies
Aditya S Patil
Abstract
Many data-driven decision policies are implicit: they sample behavior-supported actions and execute the candidate with the largest learned value or simulator score. They are easy to deploy, but their action probabilities--and hence likelihood ratios needed for off-policy calibration--are generally unavailable. We introduce candidate completion for behavior/best-of-$M$ mixtures. Given a logged action, it completes the latent candidate set, reruns the frozen symmetric selector, and applies a randomized acceptance coin. The resulting subkernel is exactly proportional to the target policy without evaluating policy densities. Composed over a finite horizon, retained trajectories are iid from the closed-loop target law at explicit overlap cost $C^{-H}$. Split-conformal ranks provide finite-sample marginal lower coverage for any frozen path target; an exact-binomial rank gives a calibration-conditional tolerance guarantee. Across three MuJoCo systems, three logger qualities, and three fitted seeds, nominal 90% bounds cover 90.78% of 13,500 fresh target-policy paths, with finite corrections in all 27 units. Validity assumes exact sampling from the known data-generating logger and does not imply statewise safety, episodic-return coverage, or policy improvement.
Chat is not available.
Successful Page Load