Verify the Signals First: The Declarative--Behavioral Gap in Metacognitive Monitoring of LLM Agents
Abstract
Verification pipelines for LLM agents increasingly use signals the agent produces about itself: e.g., confidence statements, uncertainty estimates, self-assessments of competence and conflict, etc. We argue that these signals must themselves be verified before anything built on them can be trusted. In this paper, we supply a concrete validation framework for one family of such signals: a five-dimensional Metacognitive State Vector (\msv{}) spanning emotional response (ER), correctness evaluation (CE), experiential matching (EM), conflicting information (CI), and problem importance (PI). We assess how much evidence currently supports each \msv{} dimension as a measure of its stated property using behavioural probes. While all five probes are specified here, only \CE{} is measured directly in one of its operationalizations, self-rated confidence, and evaluated herewith. We identify a central obstacle to using such signals for verification which we call the \textit{declarative--behavioural gap}, which is the possibility that what a model reports about its own state may diverge from how it actually behaves. We put forth four candidate explanations with distinct, testable predictions and we then measure that gap directly on 9 open-weight models. When asked prospectively how well their confidence would predict their correctness, all nine models described it as weakly or moderately predictive; measured discrimination over 1{,}155--1{,}188 binary trials per model was uniformly weak (AUC 0.488--0.543, mean 0.512, item-clustered bootstrap). A within-item order-reversal control showed that reversing option order changes accuracy by 0.375 on average, up to 0.92, while stated confidence shows almost no systematic shift with order, so confidence does not register the factor that governs correctness on this task. Two instruments intended to validate that signal failed in different ways, one becoming indeterminate for seven of nine models and the other returning plausible values with no detectable relationship to direct discrimination. We also propose criteria for deciding when metacognitive signals are trustworthy enough to support verification as well as an accompanying audit principle: namely, that the reported state and observed behaviour should remain separately inspectable so that the declarative--behavioural gap itself can be measured.