Time to Pay Attention! Understanding High Complexity Corpus Reasoning Tasks
Abstract
Corpus reasoning tasks, which require extracting and integrating information from large collections of documents (e.g. scientific literature, LLM agent traces), range from well-studied tasks such as factoid question answering (e.g. When was Ralph Lauren founded?), to more complex, underexplored queries like What are all the conflicting claims in this literature?. We observe that on these complex tasks, efficient architectures can scale much worse compared than they do on more well-explored tasks, but why does this happen? We develop a unified view---corpus task complexity (CTC)---to systematically distinguish complex tasks from simpler ones in terms of the smaller operations needed to solve them (and how such operations scale with corpus size): for example, we classify tasks which compare all possible document pairs as more complex than factoid retrieval. We first investigate new long-context language model (LCLM) training dynamics challenges on these tasks, and propose an effective new mask mixing technique that improves O(N) attention approaches on high CTC tasks, but at larger corpus scale we ultimately find that there is no free lunch: (i) methods that scale efficiently with corpus size do poorly when CTC is high, and (ii) methods that can handle these tasks scale quadratically in cost with corpus size. We identify high-CTC tasks as a long-term open problem and encourage more future work investigating new methods to overcome these limitations.