More Can Be Worse: Long-Context Evidence Composition under a Fixed Token Budget
Abstract
Scaling context is useful only when a model can preserve and exploit the additional information made available to it. We study this problem at a compressed long-context interface, where increasing amounts of retrieved evidence are mapped to continuous representations before generation. We find that greater evidence breadth can sharply reduce downstream utility when independently compressed representations are composed by concatenation. On a 500-example HotpotQA evaluation, two selected packets achieve 48.5 F1, while six packets fall to 39.2 F1 and using all available packets collapses to 1.6 F1, despite retaining the highest-ranked evidence. We introduce PACKETRAG, a question-conditioned composition mechanism that converts a variable-width evidence set into a fixed four-token generator interface. PACKETRAG preserves the top-ranked compressed evidence as a stable base and integrates additional context through a learned residual update. Across four QA datasets, six-packet fusion improves over independent concatenation by 2.8–10.2 F1 and over sparse selection by 2.1 F1 on average. Its performance remains comparatively stable as evidence breadth expands to twelve or all available packets. These results expose a long-context composition bottleneck: making more evidence accessible does not guarantee that its representation remains usable after composition. They motivate treating context composition—not only context-window capacity—as a first-class design and evaluation problem for long-context foundation models.