Verifier Pass Is Not Skill Use: Qualifying and Preserving Verification Contracts for Skill-Grounded Agent Data
Minhao Li ⋅ Zijian Liu ⋅ Hongwei Cheng ⋅ Huizhu Lyu ⋅ Huadong Guan
Abstract
Executable verifiers can certify an accepted final state without establishing that their reward-relevant assertions follow from the public contract or remain valid after a task is modified. We study this verification problem for long-horizon tasks with mounted external Skills. Our stack makes the verifier itself an object of qualification: controlled outcome mutations probe specified failure classes, and every reward-relevant assertion is mapped back to the public task, environment, or Skill. Separate gates then measure solver-relative Skill dependence, deterministic outcome success, and substantive timely Skill use. We instantiate the stack through Skill-Preserving Recursive Synthesis, which adds verifier-backed requirements while freezing Skill bytes, reviewed Skill scope, and truth scope; every child must re-establish the full contract rather than inherit its parent's evidence. Because the behavioral gate is model-based, we audit it with an otherwise unused model family, obtaining 86.0% agreement (Cohen's $\kappa$ 0.72) on a balanced stratified sample. The layers detect nonredundant failures, and an equal-size outcome-only control matched on teacher and training configuration indicates that substantive-use filtering changes what is learned: it reaches +0.040 on E2, leaving a +0.039 residual under the matched filtering comparison. Full-corpus SFT increased the designated primary E2 mean normalized verifier reward from 0.418 to 0.497, a paired task-level difference of +0.079. The E3 point estimate was smaller; on SkillsBench v1.1, tasks solved by at least one rollout increased from 3/87 (3.45%) to 8/87 (9.20%). Together, these results suggest that qualifying and preserving verification contracts can improve the quality of Skill-grounded training data beyond outcome verification alone.
Chat is not available.
Successful Page Load