Executable Scientific Contracts for Research Auditing
Amir Reza Peimani
Abstract
As AI systems participate in scientific writing and review, a plausible paper statement, implementation, and released artifact can still encode different scientific objects. We study executable scientific contracts, small checks that bind a stated expectation to an implementation path, discriminating fixture, released-artifact activation, and the strongest consequence supported by evidence. The contract is a rerunnable evidence object: unlike automated discrepancy discovery, it separates a local mismatch from whether the triggering condition occurs in released support and from evidence about a reported result. We instantiate 25 contracts across five recent agent and LLM research repositories: AgentPRM, ContractBench, ToolACE, $\tau$-bench, and a prospective AgentAbstain case. Among the 20 original checks, 12 were conformant, 7 discrepant, and 1 inactive; among five prospectively frozen AgentAbstain checks, 3 were conformant and 2 exposed consequential choices left unspecified by the source. All 30 study-specific mutations were detected. An external mutation set was detected in 7/8 cases on its first execution and 8/8 after correcting the missed fixture. Once authored, complete contract suites executed in under 0.4 seconds median per repository on a CPU. Scientific contracts therefore complement agentic paper–code auditing by preserving how far a detected discrepancy can actually be justified.
Chat is not available.
Successful Page Load