Recourse as a Contract: Certifying Realized Actions under Behavioral Deviation and Classifier Drift
Sita Bissu ⋅ APARNA MEHRA ⋅ Sandeep Kumar
Abstract
Machine learning systems now mediate consequential decisions in banking, healthcare, hiring, and resource allocation. When such a system returns an unfavorable decision, algorithmic recourse prescribes actionable changes an individual can make to obtain a favorable one [1]: a rejected loan applicant may be told to reduce revolving debt by a fixed percentage before reapplying. Recourse is therefore more than an explanation: it is a promise that effort will be rewarded, and whether that promise survives deployment determines individual agency and institutional trust. Existing methods, including robust ones [3], primarily certify the action the system \emph{prescribed}: that it would flip the decision of the classifier that \emph{issued} it. Deployment preserves neither the prescribed action nor the issuing classifier. The individual may comply partially, substitute a cheaper action, or respond strategically; the institution may retrain, recalibrate, or replace the model before reapplication. Robust prescribed-action recourse can thus protect against model variation while certifying an action the individual may never take: in our experiments, a robust prescribed-action method attains essentially zero prescribed failure while 10--19\% of individuals are still rejected on reapplication. We propose Recourse as a Contract (RAC), which instead certifies the deployment event itself: whether the action a recommendation actually induces is accepted by the future decision rule. RAC separates the recommendation issued by the institution from the action realized by the person, treating the human response as an unknown stochastic kernel---of which prescribed-action recourse is the degenerate identity case---so the guarantee requires no structural model of private costs or compliance. Validity comes from calibration on logged recommendation--response data under an encouragement design; a learned behavioral model may guide recommendation selection but does not enter the validity guarantee. Before issuing a recommendation, the institution either certifies it at a user-specified failure level or abstains. We characterize both the possibilities and limits of such contracts. Exact individual-level distribution-free certification is impossible, yielding the recourse analogue of known limits on conditional predictive inference [4]. We establish three complementary routes: exact finite-sample contracts on context groups via one-sided binomial tests with simultaneous validity; individual-level contracts achieving the minimax rate under smoothness; and policy-level certification from observational logs, extended with anytime-valid monitoring and off-policy certification for continuously operated systems. Certifying a recommendation whose risk lies a margin $\Delta$ below the target threshold requires a sample size growing as $\Omega(\Delta^{-2})$, making abstention near the boundary statistically unavoidable rather than an artifact of conservative calibration. A behavioral-controllability bound further caps how much realized risk any recommendation can change, identifying when recourse can be validly certified yet remain unable to materially improve outcomes. Experiments on German Credit, Adult, and FICO HELOC under controlled partial compliance, implementation noise, and classifier drift validate the theory. Across 4,000 benchmark-derived interactions, 120 drift draws, and 50 repetitions, RAC maintains calibration at requested failure levels from 0.10 to 0.30, while certification coverage falls from 100\% to 82\% as the target tightens. Required sample size scales as $\Delta^{-2}$ as predicted, with a fitted log--log slope of 2.09. As compliance degrades, RAC's realized failure rises from 1.5\% to 13.1\%, compared with 10.0\% to 18.8\% for a robust prescribed-action baseline despite its numerically zero prescribed failure.
Chat is not available.
Successful Page Load