EvoAudit: Qualifying Feedback in Self-Evolving Search under Unreliable Evaluators
Abstract
Self-evolving systems turn evaluator measurements into persistent state: parent selection, optimizer updates, archive membership, or compute allocation. When evaluators are noisy, non-stationary, or exploitable, a high score need not be safe evolutionary credit. We study feedback qualification: separating an observed development score from permission to write that observation into future search state. EvoAudit combines confirmation-based audit gates, a family-aware diversity reserve, and an information-availability contract that prevents delayed or sealed evidence from being backdated into earlier decisions. Under an equal evaluator-call budget and 50 paired seeds, EvoAudit lowers false feedback rate by 0.390 versus score-only search (95% paired-bootstrap CI [-0.485,-0.295]) and raises normalized family entropy by 0.555 versus audit-only search ([0.533,0.576]). The gain is not free: EvoAudit loses 5.677 clean feedback items per 100 calls versus audit-only search, falsifying a frozen-before-run efficiency hypothesis. A 27-condition sensitivity grid preserves this contamination-diversity-cost trade-off. We additionally report an anonymized prospective symbolic-program search case showing how sealed evidence and fail-closed commits keep invalid observations out of adaptive state.