CFR-X: Measuring What Matters in Agentic Systems
Punit Maheshwari ⋅ Siby C Pulikottil ⋅ Sriram Srinivasan ⋅ Srinivasan Rajendran ⋅ Madabattula R Kumar ⋅ Srinivasa Rao Aravilli
Abstract
While autonomous agents make continuous sequential choices, overall trajectory outcomes are typically governed by a critical minority of decisions. The challenge lies in identifying which decisions are pivotal. Taking inspiration from the game-theoretic concept of counterfactual regret (CFR), we adapt this framework to agentic systems. An agentic system is a partially observable Markov decision process, so the agent commits to a long chain of decisions from incomplete state, and the reward it eventually receives says nothing about which of those decisions actually carried it. Unlike methods that predict importance, our proposed approach instead relies on verifiable rewards, re-running alternative actions in the real environment to measure how much each decision changes the outcome. We call this Counterfactual Regret analysis for agentic systems (CFR-X). The same signal supports three capabilities usually studied in isolation. Through explanation, it ranks an agent's decisions by how much each one mattered and surfaces hidden faults. Through steering, its scores are stored and reused at inference time to redirect effort toward the decisions that matter. Through repair, it branches structurally different candidate actions and keeps the one that measurably works. We validate each capability in a different setting. For explanation, CFR-X ranks the pipeline decisions of an ML-engineering agent, finding that critical decisions outweigh the rest of the pipeline by up to $4.46\times$ and uncovering a silent metric bug that static analysis misses. For steering, a memory of past scores raises a science-world agent's task score by $43\%$ while spending $40\%$ fewer tokens on held-out tasks. For repair, branch-and-select lifts a code agent's solve rate from $54\%$ to $79\%$, and regret variance predicts which failures are recoverable ($p < 0.01$). One pattern runs through all three: importance concentrates in a handful of decisions, so knowing where an agent should stop matters as much as knowing where it should act. CFR-X applies to any agentic system whose actions can be executed and scored.
Chat is not available.
Successful Page Load