Beyond Endpoint Success: Trustworthy Completion for Web Agents
Ke Ning ⋅ Luyao Zhang
Abstract
Web agents may reach a requested endpoint while accepting an unwanted charge, granting an unnecessary permission, or disclosing personal data. Endpoint success alone therefore does not reveal whether an agent protected the user’s interests during execution. We introduce Trustworthy Completion, which independently measures nominal completion $C_r$ and avoidance of a prespecified, machine-verifiable unsafe commitment $S_r$, with $TC_r=C_r\land S_r$. We instantiate the framework in 12 synthetic consumer tasks spanning forced action, sneaking, and interface interference, and evaluate one frozen vision-capable web agent across three conditions: No safeguard, System-delivered safeguard, and Interface-delivered safeguard. All 108 scheduled cells yielded valid outcomes. Without a safeguard, the agent completed 34/36 runs (94.4%), yet only 7/36 (19.4%) were trustworthy completions; 27/36 (75.0%) were unsafe completions. Relative to No safeguard, both safeguard strategies increased observed safety by 16.7 percentage points, while System and Interface delivery reduced completion by 11.1 and 16.7 points, respectively. Trustworthy-completion gains were smaller and uncertain, and the direct comparison did not distinguish the two delivery strategies. These results expose a substantial capability–trustworthiness gap and show why safeguards must be evaluated jointly for safety and useful completion rather than through endpoint success alone. By grounding stakeholder-protecting boundaries in deterministic trajectory evidence, the framework provides an auditable and cost-aware basis for evaluating trustworthy delegated action on the Web.
Chat is not available.
Successful Page Load