Failed Proofs Are Sticky: Rejected-Proof Perseveration in Lean Recovery
Abstract
Aggregate verification records whether recovery succeeds, but not whether a scaffold revises a rejected trajectory or perseverates on it. We use perseveration behaviorally for observable reuse of verifier-rejected proof material, not as a claim about an internal cognitive state. Lean’s deterministic diagnostics let us audit that distinction, then isolate raw-attempt access in a frozen 2 × 2 experiment on 40 held-out miniF2F failures, crossing the exact previous completion with fixed failure-conditioned guidance while holding root, seed, sampler, token budget and conversation structure constant. We repeat the same theorems under a second seed and retain a 43-root development cohort for descriptive use. We define first-error source reuse (FESR): whether the source line associated with the original attempt’s first reported Lean error appears in the recovery proof. On 21 roots with trustworthy source locations, FESR occurs in 20/21 prior-conditioned proofs versus 5/21 fresh proofs in block 1, and 20/21 versus 8/21 in block 2 (paired differences +71.4 and +57.1 points; exact p = 0.000061 and p = 0.001831, exploratory). Six branches resubmit byte-for-byte an input that the same deterministic checker had already rejected. An exploratory trajectory analysis shows that the increased stability is not explained by generic long context: across seeds, tactic-sequence LCS is 0.703 for the theorem’s own rejected proof, 0.397 for fresh generation, and 0.375 when the identical wrapper carries another theorem’s length-matched failed proof. In a secondary contrast on the tested attempt-without-diagnosis interface, verification falls from 9/40 to 2/40 and from 9/40 to 1/40. No theorem is rescued uniquely by prior exposure; the lost branches generally enter Lean checking and fail during proof elaboration/checking rather than through formatting or budget exhaustion. The preregistered guidance-by-artifact interaction and the feedback-bearing attempt contrast remain unresolved. The contribution is therefore a verifier-localized account of perseveration: a raw-history scaffold can preserve first-error-associated source from a deterministically rejected input, so continuity must not be mistaken for correction.