Probing the Trajectories of Reasoning Traces in Large Language Models
Abstract
Large language models (LLMs) increasingly solve difficult problems by producing reasoning traces before emitting a final response. However, it remains unclear how accuracy and decision commitment evolve along a reasoning trajectory, and whether intermediate trace segments provide answer-relevant information beyond generic length or stylistic effects. Here, we propose a trajectory-probing protocol for evaluating how a reasoning model's predicted answer evolves along a partial reasoning trace. The protocol 1) generates a model's full reasoning trace, 2) truncates it at fixed token-percentiles, and 3) injects each partial trace back into the model, measuring the model's induced answer distribution. We apply the protocol to five open-source reasoning models (Qwen3-4B/-8B/-14B and gpt-oss-20b/-120b) across three benchmarks (GPQA Diamond, MMLU-Pro, and Omni-MATH-2). The protocol reveals that accuracy and decision commitment consistently increase as the percentage of provided reasoning tokens grows, and length, form, and token-identity controls confirm that gains stem from instance-specific semantic content rather than context length or generic reasoning style effects. Furthermore, weak-to-strong cross-model probing experiments reveal when stronger models recover from incorrect partial traces and when they instead anchor to them. Together, these results provide a reproducible, model-agnostic measurement of reasoning-trace dynamics that generalizes across answer formats and can reveal model-specific failure modes invisible to aggregate accuracy.