When Should an Agent System Change Itself? Evidence-Gated Architecture Updates Under Non-Stationarity
Ali A Alzahrani
Abstract
Agent development increasingly changes prompts, workers, communication, model assignments, budgets and verification logic, but a generated system change should not replace a functioning incumbent merely because it scores well on noisy evidence. We study verification of agent-system updates under non-stationarity and introduce Conservative Meta-Agent Search (CMAS), a meta-agent loop that proposes bounded typed architecture edits and requires candidate-specific evidence before promotion. Throughout, meta-agent denotes the whole outer loop (proposer, compiler, evaluator, gate and rollback controller), not the proposal model alone. CMAS applies hard eligibility checks, compares candidate and incumbent by paired shadow replay under common random numbers, assigns interventional trace-level credit, and promotes only when an empirically calibrated score clears a pre-specified margin; promotion then enters a simulated canary stage with rollback. We evaluate on MetaInvest-Bench, which combines a controlled simulator with evaluator-only latent ground truth, leakage-controlled market replay, and an adversarial governance suite. Across 96 controlled streams, CMAS reduces normalized mean dynamic-oracle regret from $0.116$ to $0.104$ and raises held-out utility from $0.44$ to $0.48$ relative to the strongest adaptive common-budget reimplementation. False promotion is 3.4% of update decisions and 4.8% in a sealed rerun; removing paired evaluation or simultaneous correction raises it to 6.1% or 8.4%. Across 576 simulated governance incidents, CMAS recovers within three updates in 78% of episodes with rollback precision/recall $0.85/0.80$. A return-only evaluator produces a growing downstream utility gap, exposing objective misspecification despite successful optimization of the measured signal. The verification claims are deliberately narrow: the implemented score is empirically calibrated rather than a confidence bound and does not establish full utility, absence of canary harm, sequence-level error control, or live-deployment safety.
Chat is not available.
Successful Page Load