GeoReason: Step-Level Hallucination Detection via Hidden-State Trajectory Geometry
Abstract
Large language models hallucinate during multi-step reasoning, but most existing detectors operate at the trace level: they assign one confidence score to a full output, fail to localize the first error, and often require multiple sampled completions. We frame hallucination instead as a property of the hidden-state trajectory produced during a single forward pass: correct reasoning moves through a stable manifold of locally coherent transitions, and a first error appears as a localized excursion in transport cost away from this manifold. We investigate this hypothesis with a label-conditioned teacher that builds a trace-specific contrastive PCA representation and measures geometric trajectory dynamics. Although not deployable, it serves as a diagnostic instrument for identifying transferable geometric signatures of reasoning failure. We prove that contrastive PCA is the optimal projection for a transport-separation objective between first-error and correct states, and that single-pass first-error localization holds whenever the first error creates a positive transport margin over preceding correct transitions. Across ProcessBench, PRM800K, HaluEval, and TruthfulQA, the teacher consistently outperforms entropy-based, probing-based, and attention-based baselines while retaining substantial signal across models and datasets. We further explore deployment through a distilled BiLSTM operating on raw hidden states. While this student performs well in-domain, it degrades under distribution shift, suggesting that recovering transferable geometric dynamics is substantially more difficult than detecting them.