Reading Is Not Using: A Retrieval--Integration Gap in Long-Context Decision Making
Abstract
Long-context evaluation is dominated by recall: can a model locate and restate a fact in a long input? Prior work shows, in aggregate, that task performance can fall with length even when retrieval succeeds; we measure that gap per passage. On a financial decision task, holding decision-relevant content byte-identical and growing only unrelated text from 2K to 128K tokens, we estimate how much one passage moves the model's judgment against matched neutral replacements and a placebo-insertion null. For the primary model that influence falls to the null between 8K and 32K tokens while directed recall of the passage stays at ceiling; the decline replicates across three open-weight families and is echoed, under a complementary readout, in a production system and in an exploratory arm that redacts genuine disclosures from real SEC filings. Because matched documents differ in one slot, the model's routes can be cut and spliced: at short context, attention and a hybrid model's recurrent state each carry a large, statistically indistinguishable share of the passage's influence, and with length the passage stays encoded where it is read while content-selective attention to it collapses, evidence that favours a failure of transmission over one of reading. Extended reasoning does not restore influence and a generic chunk-and-summarize pipeline evicts the passage at extraction, whereas a targeted structured restatement adjacent to the decision raises retention at 128K from 12% to 67%, which neither proximity nor a verbatim self-copy of the passage reproduces. Recall at a given length does not certify use at that length; per-passage influence, judged against a placebo null, measures it.