When Verification Changes the Verdict: A Three-State Audit of Active GUI-Agent Evaluation
Abstract
Active GUI-agent evaluators inspect live applications because final screenshots hide persistent state. But an inspection can also change the success predicate it is supposed to verify. We introduce a three-state audit that freezes actor-end truth before inspection and separately records the verifier's working state and the original source. Across 360 independently initialized trajectories in a frozen 24-state testbed, Active-Original flips the task predicate in 54.2% of trajectories and yields actor-end label errors in 51.4%. The sharpest result is a separation of guarantees: Transactional execution reduces source mutation to zero, yet 47.2% of Transactional trajectories yield actor-end label errors aligned with a verifier-created fork state. A disposable fork can preserve the source environment while the verifier credits the actor for a state created inside that fork. Hard-ReadOnly eliminates target shifts on this testbed but sacrifices hidden-state coverage. Two native applications locate the boundary of the failure chain: Roundcube does not elicit inspection, whereas Excalidraw inspection flips all three predicates but preserves correct actor-end verdicts. Thus mutation is a genuine measurement hazard, not an inevitable label error. Our protocol makes probe selection, predicate change, source integrity, and label validity separately measurable.