Active Oracle Evaluation for Stochastic Decision Agents: A Backgammon Case Study
Nikash Bhardwaj ⋅ Shantanu Jaiswal ⋅ Deepak Pathak
Abstract
Outcome-based evaluation fails when noise exceeds the skill gap being measured. In Backgammon, the strongest agents differ by a fraction of a point of expected score per decision, a gap dice variance buries in whole-game results, leaving win rate and Elo nearly uninformative. Recasting evaluation as decision-quality measurement, we treat GNU Backgammon as an engine oracle estimating the equity of every legal move. Per-decision regret against it rates each agent, separating a pair with $8.1\times$ fewer trials than points-per-game needs at equal precision. Since labels are costly we spend a fixed budget on decisions where agents disagree or best separate near-tied pairs, recovering the ranking from half the labels; a Horvitz–Thompson estimator keeps ratings unbiased under any selection rule. The rating reproduces the ordering of independent head-to-head play (Kendall's $\tau = 0.93$) and calibrates to win probability (R$^2 = 0.95$). Value-based, policy-gradient and search-trained networks cluster tightly, a classical public-domain evaluator sits inside that cluster, and AlphaZero's advantage is flat across game phases rather than concentrated in the contact positions where search helps most: the signature of a stronger evaluation function, not deeper search. The protocol's conditions are met well outside games, and so is its scope condition: closed-loop driving evaluation reports the same near-parity inversion. Code available at: https://github.com/Anonymous-376/active-oracle-evaluation.
Chat is not available.
Successful Page Load