Reasoning Pathologies in Large Language Models: A Diagnostic Perspective
Abstract
Long chain-of-thought (CoT) reasoning in large language models often fails in recurring ways: models loop through redundant derivations, drift through uncertain intermediate states, commit prematurely to unsupported answers, or continue generating after a valid conclusion has been reached. These failures are typically treated as separate surface phenomena and addressed with symptom-specific heuristics, which (i) leave the underlying trajectory failure unchanged when they suppress the visible symptom, and (ii) provide no shared substrate for comparing fixes across models or pathologies. We instead study them as failures at the level of CoT reasoning latent state transitions. We introduce a diagnostic framework that abstracts chain-of-thought reasoning as a trajectory through discrete latent reasoning states, yielding a positional taxonomy of six reasoning pathologies, each paired with a transition-grounded detection predicate and a matched inference-time controller. We evaluate the framework through controlled-surgery experiments across multiple datasets and models, spanning arithmetic reasoning, multi-hop question answering, and knowledge-intensive reasoning. Across evaluations, we find that reasoning pathologies exhibit three main diagnostic patterns: \textbf{(1)} failures follow a phase-ordered structure, where termination failures are most reliably identified, early branching failures are detectable, and progression failures require finer graph-aware analyses; \textbf{(2)} counterfactual transition-versus-emission tests show that several detectors track latent trajectory changes rather than surface rewrites, while also exposing cases where paraphrases alter the reasoning path itself; and \textbf{(3)} detector-timed interventions are most effective when the detected pathology has a clear local onset, but do not uniformly dominate generic controls. Together, these findings position the framework as a reproducible diagnostic substrate for studying CoT failures, with explicit evidence boundaries and stronger closed-loop control.