Reading Between the Dots: Decoding Hidden Computation across Filler Tokens
Abstract
Frontier LLMs can perform multi-step reasoning over content-free filler tokens like dots or counting sequences, producing correct answers with no visible chain-of-thought (CoT). This has been cited as a limit case for CoT monitorability: if the surface tokens carry no information about the reasoning, behavioral oversight has nothing to read. But hidden from the output is not the same as hidden from us. On three task families spanning fact retrieval, parallel numeric composition, and string manipulation, two open-weights frontier models (DeepSeek V3, Kimi K2) compute over filler tokens in a structured, legible way. Attention routes the question through the filler region to the answer, KV-cache transplants at only filler positions causally swap the model's output between examples, and logit-lens readouts reveal a temporal structure in which retrieved facts appear early and their composition crystallizes in late layers. We then introduce an unsupervised decoding pipeline that takes only hidden states as input and recovers intermediate values with 80–95\% accuracy with the strongest LLM judge across both models and all three tasks, without ground-truth labels or training. Hidden computation that defeats behavioral CoT monitoring is, on these tasks, directly readable from the residual stream. This suggests that monitorability should be understood as a property of the model's full computational trace, not just its surface tokens.