BioMaster Review: Auditable Integrity Screening and Reproduction Contracts for AI-Native Peer Review
Abstract
Scientific peer review increasingly requires verification across heterogeneous evidence: numerical values, table relations, source artifacts, figures, code, and computational results. However, existing tools typically expose one signal or begin from a predefined reproduction task, leaving both evidence aggregation and task construction implicit. We present BioMaster Review, an AI-native, evidence-centered peer-review system implemented as a reusable paper-review skill. Tier 1 defines a 15-check catalog in four parallel integrity layers and executes applicable checks while retaining applicability, raw statistics, thresholds, and evidence locators before aggregation. A review-priority layer routes selected claims to Tier 2. There, QAminer defines a paper-to-reproduction-contract protocol that separates claim, method, and data extraction from implementation and post-implementation acceptance-criteria mining. The system returns inspectable evidence for human review. We evaluate progressively deeper operations on three benchmarks and one paper case. On CPTAC numerical fabrication, controlled models reproduce published performance, while detector traces expose failure of an uncalibrated rule vote. On FakeParaEgg, failure-aware routing attains macro F1/IoU of .477/.394, comparable to the strongest published out-of-distribution point estimate (.478/.396), with fixed fusion yielding lower scores under the same evaluation. On ClaroAI-Bench, the separately executed workflow obtained 50.9% overall versus 49.4% for the human-curated reference, with positive D5 matching on 20/33 versus 18/33 papers and exact per-paper agreement on 24/33 (72.7%). The resulting evidence contract provides a concrete basis for connecting screening, routing, and reproducible computation as separately auditable review operations.