What Should a Guardrail Return? Factorizing Post-Block Feedback for Browser Agents
Usukhbayar Purevdorj ⋅ Zekun Wu ⋅ Adriano Koshiyama ⋅ Philip Treleaven
Abstract
Browser agents can take consequential actions on websites, so a safety check can inspect a proposed action against what the page really does before it runs. Blocking an unsafe action prevents harm but can also leave the user’s task unfinished. We keep the initial blocked condition identical across runs and vary two pieces of feedback, each a fact the safety check itself verified. Cause feedback says why the action was blocked. Availability feedback says a permitted way to finish the task still exists. Five fixed feedback conditions produce 1,440 runs on one classified-ads site: 18 product listings, four models, two repeats, and two ways of showing the page, as structured element text or as a screenshot. Recovery means permitted task completion after the unsafe action is blocked. Availability feedback reduced recovery by 8.13 percentage points on average (95% interval $[-10.45,-5.80]$; exact $p=3.05\times10^{-5}$). Cause feedback improved recovery by 10.93 points ($[8.75,13.00]$; Holm-adjusted $p=3.05\times10^{-5}$). The interaction was +7.37 points ($[2.18,12.74]$; Holm-adjusted $p=.0156$): the availability penalty shrinks when cause is present. Forty-one runs returned an empty model response we could not score, all from one model, GPT-5.6 Sol, on screenshots. Even if all 41 are scored in the way least favourable to us, the three effects keep their signs. No prohibited action reached the browser. In this setting, explaining why an action was blocked was more useful than simply saying that another valid action existed. One testable reading is a shift toward one-shot alternative selections, with cause narrowing the choice, under an endpoint that makes a wrong alternative terminal; the present experiment does not establish this mechanism.
Chat is not available.
Successful Page Load