Reading the Plan: Multi-Horizon Future-Action Prediction from Frozen Hidden States in LLM Agents
Laksh Advani ⋅ Suraj Maharjan ⋅ ISHA CHATURVEDI
Abstract
Tool-using LLM agents operate sequentially: they reason, call a tool, observe its result, and repeat. Each call blocks on execution, leaving the accelerator idle with no cheap signal about what comes next. We show that signal already exists: future tool calls are linearly decodable from a single frozen residual-stream state at roughly three-quarters model depth, a byproduct of the forward pass that commits to the current action. A 1.01M-parameter MLP reading this state predicts the tool call ($k=1,3,5$ steps ahead) at 0.09 ms per decision, requiring neither fine-tuning nor a second model copy. Across three architectures spanning a 3.4x parameter range, $k=1$ accuracy varies by under two points. Against a matched one-hot-only classifier, the hidden state adds 20 to 35 points at $k=1$, widening with horizon, ruling out current-action identity as the source. On a 110-tool benchmark, the probe matches full-corpus Markov counts overall at one step ($k=1$) but nearly doubles their accuracy on hard, low-confidence examples (0.562 versus 0.306) and surpasses them under confidence gating. It also outperforms the model’s verbalized predictions by 12–30 points even under the three shot chain-of-thought, showing that future behavior is encoded in activations but poorly accessible through self-report. On smaller vocabularies, the probe matches or beats QLoRA at a fraction of the cost. A meaningful share of an agent's future behavior is already sitting in its activations, cheaply readable for prefetching, scheduling, and monitoring.
Chat is not available.
Successful Page Load