Investigation Capability versus Investigation Intelligence: A Controlled Study of Selective Belief Updating
Yashash Yallapragada ⋅ Shambhavi Srivastava
Abstract
An agent with persistent beliefs must decide whether an unexpected failure justifies changing them: UPDATE risks learning from noise, IGNORE risks missing genuine change. We study a third action, INVESTIGATE, that pays a cost to acquire evidence before committing, and build a controlled benchmark where every failure's true cause is known but not revealed by its surface description. Positive result: non- investigating policies sit on a stability-adaptability tradeoff curve (perfect stability or reasonable adaptability, never both); every investigation-capable policy we test escapes this curve. Negative result: across a budget-matched control, a dense budget sweep ($B\in\{0,1,2,3,5,8,\infty\}$), two harder ambiguity regimes, and an evidence-information ablation, an LLM gate that reasons over each failure and its own history to decide when to investigate shows no statistically significant advantage over simple heuristic or budget-matched random triggers (11 of 12 paired comparisons non-significant; the one nominal exception does not survive correction for multiple comparisons). Counterfactual and diagnostic-regret analysis finds no evidence, within this benchmark, that downstream evidence acquisition or interpretation explains the residual gap between the LLM gate and an oracle ceiling; the residual gap is most strongly associated with the initial trigger decision. Diagnostic investigation is valuable; the intelligence of deciding when to invoke it is not demonstrated by the LLM policies we test.
Chat is not available.
Successful Page Load