Counterfactual Evaluation of Agents in Non-Resettable Environments
Abstract
Deployed language agents take actions that cannot be undone. They send email,move money and write to production databases. The ordinary way to evaluate a proposed change is to ship it and watch what happens, which is precisely what the operator of such a system cannot afford. Offline evaluation ought to fill the gap, but off-policy evaluation (OPE) assumes repeatable transitions and small action spaces, and agent deployments supply neither. Token-level importance weights collapse in effective sample size, and scaffold changes that add tool calls break absolute continuity. We present ARC-DR, which scores candidate agent configurations against historical logs without executing anything irreversible. It leans on two properties of agent systems. Irreversibility attaches to individual actions rather than to whole trajectories, so every episode has a commit point, the index of its first irreversible action, before which a local snapshot permits exact re-execution. And although the environment cannot be restored, the policy remains available for querying, so propensities can be estimated by replaying logged context windows rather than by fitting a density model over text. ARC-DR re-executes the candidate live up to its counterfactual commit point, intercepts it there, and corrects the remaining suffix from logs using doubly-robust per-decision estimates over a semantic action abstraction. We prove that variance falls exponentially in the commit index, bound the bias contributed by abstraction leakage, suffix matching and a finite replay budget k, and give sharp Manski bounds for candidates that expand the action support. Over m= 24 variants on a stateful benchmark ARC- DR reaches rank correlation rs = 0.83 (95% CI [0.64,0.92]) at MAE = 0.036, beating a live 10% budget baseline on point accuracy while that baseline spends 50 side-effecting executions per candidate. In a sandboxed deployment against a live database, a payment API and an SMTP sender, interception held for all 84 irreversible calls. It did not hold under adversarial probing. A static tool allowlist let 14 of 20 implicit commits through, and only a conservative fallback that treats an ambiguous tool as committed at cā²= 0 caught them all. Classifying side-effects inside dynamically generated tool calls is still open.