Aggregate Accuracy Is the Wrong Quantity: A Due-Process Standard for Detector-Assisted Desk Rejection
Yuvanguru Balagurumoorthy
Abstract
Program committees are converging on detector-assisted triage: classifiers that score submissions for AI generation, with high-scoring papers routed to desk rejection or integrity review. This is happening precisely as the detection literature documents nontrivial error rates and demonstrated paraphrase-based evasion. We argue that the quantity venues cite when justifying such systems — aggregate accuracy or AUC on a benchmark — is the wrong quantity for a decision that ends a specific author's submission. The right quantities are the expected number of falsely accused authors and the positive predictive value of a flag at the venue's actual base rate. We develop this parametrically: at $N$ submissions, false-positive rate $p$, and true prevalence $\pi$, the expected count of falsely flagged authors is $N(1-\pi)p$, and PPV collapses toward zero as $\pi$ falls, however impressive the ROC curve. At conference scale, single-digit FPRs yield hundreds of false accusations per cycle. We then argue an asymmetry: the authors most likely to be falsely flagged — non-native English writers, whose prose is statistically typical — are not the authors most able to evade detection, since evasion requires only adversarial paraphrasing. Detection thus concentrates its costs on the population least able to avoid them and least equipped to appeal. We close with a six-clause minimum due-process standard a program chair can adopt directly into a call for papers, and argue it is compatible with, not opposed to, scaling review past 30,000 submissions.
Chat is not available.
Successful Page Load