What Carries the Calibration Signal in Incremental Question Answering? A Pre-Specified Confirmatory Decomposition
Abstract
Incremental question answering gives a system progressively more of a question. At every step it must choose between answering now and waiting for more text. We ask whether confidence improves when the estimator uses the answerer’s recent history, not only its current score. The intuition is simple: an answer that has persisted while its score rises may be safer than an answer that just appeared. We develop this idea on a development split, freeze eight directional hypotheses, and evaluate once on an untouched test split with 63,050 prefixes from 1,953 questions. A full model lowers Brier score by 0.02253 relative to a linear base-feature estimator. Decomposing that number changes the interpretation: 29% comes from information that an offline dataset reveals but a deployed system cannot know yet, 23% comes from changing the estimator class, and 47% comes from trajectory features available at decision time. Score movement contributes more than answer stability on the confirmatory split, but that ordering does not remain distinguishable on an adversarial split. The complete causal bundle gives no detectable Brier improvement for the tested generative answerer. At the decision level, simpler base-feature estimators capture the most reliable gain. We also show how merging unresolved generated answers into one identity inflated our own stability estimate about ninefold. The paper is therefore an example of why an attractive feature story should be decomposed before it is named or generalized.