Early Hallucination Detection in Small Language Models via Confidence Trajectories
Abstract
Small language models are attractive for low-latency and cost-sensitive applications, but their susceptibility to factual hallucinations makes efficient hallucination detection especially important. Practical hallucination detectors for these models must operate under the same resource constraints. We study whether the generating model’s token-level confidence profile can provide a compact white-box signal for detecting factual hallucinations, including before a response is complete. We present Trajectory of Evolution of Confidence across Tokens (TrajECT), a method that uses a compact feature set while remaining competitive with richer hallucination detection methods. TrajECT represents each response as the tokenwise evolution of entropy, maximum probability, and entropy change, then scores the sequence with a small recurrent model to predict whether the response contains a factual hallucination. Across closed-book responses from six question-answering datasets, its leave-one-dataset-out area under the receiver operating characteristic curve (AUROC) is competitive with feature-rich white-box hallucination detectors. Across two hallucination failure modes, TrajECT is also a top performer in cross-failure-mode transfer among learned detectors. These results identify confidence trajectories as a compact, cross-dataset signal for early factual hallucination detection in small language models. We publicly release the code and evaluation pipeline for TrajECT.