When the Gates Fire: Machine-Screened Item Construction for a Discourse-Restoration Evaluation
Abstract
Definite null instantiation (DNI), a frame-semantic argument that a sentence omits but that prior discourse makes recoverable, is a natural substrate for testing whether foundation models exploit discourse structure. We turned every usable DNI annotation in FrameNet 1.7 full-text into 1,863 machine-screened four-option items probing whether restoring the omitted argument improves model accuracy, and we validated the items before evaluating any model on them. Four machine gates (adversarial leak probes, twin-answerability, structured-evidence linguistic validation with verified source quotes, and an aggregate battery verdict) preceded a preregistered human gate requiring zero critical validator false-passes and close agreement between two independent experts. Both preregistered gates fired: the machine validator false-passed invalid items under both the original and a later clarified reference, and the experts disagreed beyond threshold, with post-hoc adjudication revealing a directional calibration difference in the first expert's labels rather than rubric ambiguity. The confirmatory attempt therefore stopped before evaluatee calls, and a preregistered three-candidate validator-redesign tournament also produced no eligible winner. Operationalizing the linguistic category itself proved unstable: a bright-line rule separating discourse-stated from world-knowledge-recoverable antecedents moved 15 of 30 expert DNI-validity labels. On a mixed-certification 30-item set (7 human-anchored, 23 agent-screened), a separately governed three-arm descriptive probe found the restored antecedent beating a length-matched same-document placebo by +13.3pp (bootstrap 95% CI [+3.3, +23.3]), while neither no-context contrast separates from zero; a skewed key mix makes absolute levels uninterpretable, so we read only within-item contrasts. In plain terms: the true antecedent reliably beats a same-document distractor, but the probe cannot yet distinguish restoration benefit from placebo distraction.