Admissible, Not Instructive: Under-determination in Self-Play Post-Training Rewards
Abstract
Self-play post-training promises improvement without human data: a model proposes its own tasks, a verifier grades them, and a learnability reward pays the proposer for how often the solver fails. We show this objective is under-determined. Writing φ(τ) for the probability the solver fails on task τ, the reward is a function of φ alone, so two tasks with equal failure probability are worth the same however much they differ in what they teach, and reshaping the reward cannot separate them. We measure the consequences in a 0.5B Absolute Zero Reasoner whose only reward signal is a Python interpreter. The scalar such systems log and optimise against confounds solver skill with task difficulty; a decomposition built from checkpoints a run already writes separates the two, and finds the difficulty marginal moving nine times as far as the skill marginal, which does not move at all. Instrumenting the reward by category shows the proposer's gain is the policy learning to stop being rejected, and that none of it comes from the payout on proposals admitted and scored, whose movement offsets part of the gain instead. Two task sets matched on solve rate to a verified residual gap of 0.011, so the reward values them to within that, differ by 0.143 in accuracy gain, unchanged by tightening the match sevenfold across three independently built pairs. Reshaping the reward does not reach this. Only the verifier, which sees the task rather than an outcome, closes the axis it names, and it does so on both seeds; the two corrections we tried compose, and neither is sufficient alone.