Models Quote, Code Verifies, Referees Judge: When the Verifier Is the Weak Link
Tzu-Chi Yen
Abstract
Referees now write with language models, and venues answer by forbidding it or detecting it afterward. We tried a third answer: trust a model exactly where a machine can check it. In $\texttt{refereekit}$ the referee decides the recommendation, and the model drafts the review from it and from words that deterministic code has matched to the manuscript. On $16$ recent arXiv papers we compared these drafts with a language model asked to write the whole review. One frontier model almost never fabricated evidence: content the manuscript does not contain is at most $7.6$% of what the free-form review cites and $0.2$% of what the drafts cite. We had built the checker to catch the model, and three quarters of the failures it reported were its own. Deterministic code, too, adds claims of its own, such as a page the model never named. When most of a verifier’s alarms are false, the referee stops reading them and misses the true ones. A verifier’s verdicts must say how much each failure matters, and it must never guess a page. A rule enforced in a tool holds whether or not the model obeys, and our checker lets a venue audit any review without a model.
Chat is not available.
Successful Page Load