What Does a Machine-Written Submission Actually Contain? A Pre-Registered Citation and Readiness Audit of an Autonomous Research Agent
Abstract
Program chairs are already running citation audits, detector pilots and adversarial-submission checks against machine-assisted manuscripts, largely without a shared protocol and largely without ground truth about what such a manuscript actually contains. We contribute a pre-registered, artifact-preserving audit protocol for exactly that question, and apply it to Sakana AI's AI Scientist v2 (v2), a system that produces complete submissions end to end. Eleven biomedical research prompts were registered on OSF before the system's experimentation phase ran. We then scored 24 v2 manuscripts and 20 v2 ideation records under a two-track workflow –an AI-drafted first pass and single-reviewer human adjudication, both preserved – against 3 peer-reviewed high-school-first-author Journal of Emerging Investigators papers as positive controls and 3 v1 template outputs. Three results bear on submission integrity. 1. Fabricated citations are measurable and non-trivial: mean 1.12 hallucinated references per v2 manuscript (bootstrap 95% CI [0.79, 1.50]) against 0 for the human-authored controls. 2. The system's own novelty signal is systematically looser than the literature supports: all 20 ideation outputs were self-flagged novel, and independent human review classified all 20 as only partially novel – an instance of AI-shaped novelty that no submission-side disclosure field would surface. 3. None of the 24 outputs met our predefined peer-review-readiness criterion, while the 3 peer-reviewed controls met 3 of 3, so the gap is not a scoring artifact. Reading-level dimensions are descriptive-similar to the controls, and craft-level dimensions (citation honesty, figure quality, methods reproducibility) are 1–1.5 points lower on a 5-point scale – precisely the failure profile that a fluency-trained human reviewer is worst placed to catch. Two silent-failure bugs fire in all 24 runs at the pinned SHA while the pipeline still exits 0 and emits a finished-looking PDF; both were disclosed to the developers via a public issue before submission. We release the protocol, the rubric, both scoring tracks and every artifact, so that the audit is a reusable instrument rather than a one-off result.