Diagnosing the Test Before Anyone Takes It: Auditing Agent-Authored ML Problems for Validity and Discrimination
Abstract
A solution can be scored—there is ground truth and a leaderboard; a definition cannot. We study whether model-authored machine-learning problem definitions can be audited before any solver runs. We had a frontier model author 32 problems from public data; on review, the failures collapse into two classes: data leakage (12 of 32, confirmed) and zero discrimination (2 of 32). A leakage audit treats every clause of the problem's contract as falsifiable and attempts to falsify each: against 54 injected single-point defects it reaches 72.2% localization with a strict false-report rate of 11.1%, and no flag on any problem known to be clean—one flag uncovered a defect we then verified against source records. Across four auditor models, misses concentrate identically on defects whose evidence spans multiple rows—the one protection the contract lacks a clause for; adding that clause lifts cross-row localization from 2/9 to 8/9. A discrimination audit fields five reference submissions on the real evaluation segment and reads only relations between them: 85.7% localization, zero flags on clean configurations, and a structural finding—only the range and the noise are properties of the evaluation itself; the contender's position belongs to the problem-and-learner.