PageGuide: Grounded, Verifiable Web Interaction
Abstract
Web agents now answer questions about pages and carry out multi-step tasks on users’ behalf, but they report what they concluded without showing where on the page the conclusion came from, or which evidence justified each intermediate step. Users are left either to trust the agent’s behavior blindly or to re-do the work to check it. We present PAGEGUIDE, a browser extension that makes a web agent’s outputs and trajectory inspectable by grounding them in page evidence: HTML DOM elements, text spans, and visually rendered content such as maps, images, and charts. PAGEGUIDE offers two interaction modes: FIND highlights the supporting evidence for an answer in situ on the page, and GUIDE walks through a navigation task one step at a time, capturing textual or annotated visual evidence at each step. In a within-subject study (N=92) in which participants judged whether an agent’s answers and trajectories were correct, grounding improved verification accuracy from 65% to 77% (p<0.01) and reduced verification time from 73s to 67s (p<0.01), while significantly reducing manual scrolling and text selection. Accuracy gains held for both correct and incorrect agent outputs and for both text and visual evidence. We discuss what these results imply for designing agents whose behavior humans can interpret at runtime.