ScholarShield: A Contamination Firewall for Human-AI Co-Hallucination in the Scientific Corpus
Abstract
A hallucinated scientific claim is usually treated as a local error in one manuscript or review. In AI-native academia, the larger risk is propagation: an unsupported claim can pass from author to reviewer, meta-review, citation, index, summary system, and eventually future training corpora. Once repeated, the claim may acquire apparent legitimacy even if no role independently verified it. We call this process scholarly contamination and propose ScholarShield, a lifecycle firewall for human-AI co-hallucination. The framework represents scholarly claims as provenance-bearing objects with explicit states: verified, unverified, contested, corrected, or withdrawn. Gates are placed at submission, AI-assisted review, citation, publication, indexing, and corpus ingestion. The goal is not to block AI-generated text, but to prevent unverifiable claims, hidden prompt instructions, and unsupported citations from silently changing trust state as they move across roles. We define a contamination graph and metrics including propagation depth, contamination amplification, provenance retention, correction latency, and closure of downstream corrections. The design directly addresses two emerging failure classes: prompt injection in manuscripts that can manipulate AI-assisted review, and recursive ingestion of model-generated content that can degrade future models. We outline a red-team benchmark for testing venue infrastructure and argue for a quarantine-and-correction model rather than detector-based punishment. AI-native academia needs not only content generation and review tools, but a mechanism that keeps uncertain claims from becoming institutional facts by repetition.