Why Action Verification Is Not Enough: Lifecycle Authorization for Stale Commitment Execution in LLM Agents
Lihe Ye
Abstract
Long-horizon agents may execute an intention after a later cancellation, rescheduling, or override has rendered it obsolete. We call this stale- commitment execution: an agent recognizes action evidence yet acts incorrectly because the authorizing commitment is no longer valid. This identifies lifecycle authorization as a missing boundary upstream of action-time verification, determining whether a commitment remains eligible. Using StagePM as a controlled experimental instrument, we intervene on the two boundaries independently. In a 30-week factorial limited to filtering backbone proposals, adding authorization to evidence-only verification reduces $3\mathrm{FP}+ \mathrm{FN}$ loss by 10.44% (26/30 weeks); the mean paired benefit is 6.73 loss units per week (95% CI $[3.53,9.63]$). The current evidence gate raises precision but does not further reduce loss after authorization, showing that separately useful verification gates do not necessarily improve end-to-end performance when combined under their present implementations. On 24 external tasks, authorization reduces loss from 56 to 8 while preserving control execution, extending the boundary effect to an independent authorization catalog. Disjoint lifecycle interventions, failure analysis, and a two-model runtime replication further support the result. We conclude that action- evidence verification cannot substitute for verifying the commitment state that authorizes action.
Chat is not available.
Successful Page Load