Stress-Testing Language Model Harnesses for Long-Context Reasoning
Abstract
Language-model (LM) harnesses enable LMs to operate effectively over long contexts using additional compute. However, existing long-context evaluations are insufficient for distinguishing modern harnesses, leading to saturated accuracy and limited separation in efficiency. In this paper, we introduce a benchmark for evaluating both the effectiveness and efficiency of long-context harnesses. Our tasks require strategic and adaptive reasoning over global and local context, where much of the context is semantically relevant but only a small subset is useful at each step, creating both a challenging search problem and different accuracy--cost tradeoffs across processing strategies. For example, one task requires identifying and reusing prior computations among hundreds of relevant code cells, rather than expensively simulating them from scratch. We evaluate multiple families of frontier language models with four state-of-the-art harnesses. Our benchmarks remain challenging even for strong model--harness combinations. E.g., GLM-5.3 with RLM achieves only 43\% accuracy. More importantly, we find that the same underlying model can exhibit markedly different efficiency under different harnesses. Our results establish efficiency as an important axis for long-context evaluation and provide a testbed for developing harnesses that process context strategically rather than exhaustively.