CascadeClarify: Verifying Clarification Policies through Downstream Violation Reduction
Abstract
Clarification policies for agents are often evaluated using local proxies such as uncertainty reduction or question relevance, even when a clarification affects several dependent outputs. We formulate one-question clarification as an environment-grounded verification problem: selecting the question whose answer most reduces severity-weighted downstream violations. CascadeClarify estimates this value by combining uncertainty with graph-mediated influence, severity, effective preventability, and question cost. We also introduce CascadeClarifyBench, a controlled five-family stress test in which these signals agree or conflict and candidate questions are evaluated under a shared latent task. On a frozen 120-case test set, CascadeClarify achieves 66.7% agreement with the realized best question, compared with 51.7% for an algebraically matched source-only ablation, while reducing normalized regret from 0.163 to 0.081. Propagation changes only 20 selections but improves realized value in all 20, indicating a targeted rather than universal benefit. Family-wise diagnostics also expose a calibration failure: multiplicative uncertainty can dominate broader downstream reach. These results establish downstream violation reduction as a verifiable outcome for analyzing clarification policies and the regimes in which their local value proxies fail.