A Validation Contract for Anticipatory Peer-Review Benchmarks
Abstract
Tools that anticipate peer-review concerns could help authors decide what to revise before submission or another review round. Evaluating such tools is difficult because success may mean matching one realized panel, performing well across possible panels, or improving a paper after an author follows the advice. We separate these targets formally and provide a prospective validation contract covering audit samples, adjudication, error metrics, capacity sensitivity, deterministic controls, and pass/fail rules. The motivating RevPlan-Bench pipeline produced thousands of canonicalized review issues and revision plans, but the underlying corpus and scoring artifacts do not survive in the project. Its assignment rates, model rankings, and inferred issue cascades consequently depend on unaudited extraction, a restrictive unmeasured pre-filter, and unrestricted one-to-many matching. We report them only as pipeline outputs. Because the available record satisfies none of the contract's validation requirements, it supports an estimand and a benchmark governance contribution, but no claim about review dynamics, model foresight, or author utility.