Recursion as Window Relief: A Controlled 2×2 Factorial of Recursive Language Models and Prompt Compression
Evandro Franco ⋅ Akarsha Sehwag
Abstract
Prompt compression and Recursive Language Models (RLMs) are competing answers to a long context, yet practitioners compose them on the intuition that a shorter document makes every recursive sub-query cheaper. We test that intuition in a fully crossed 2$\times$2 factorial, \{direct call, RLM\} $\times$ \{no compression, LLMLingua-2 at rate 0.5\}, over 200 fixed LongBench-v2 questions on seven frontier models, 5,600 runs, with exact-binomial McNemar tests and question-level cluster bootstraps throughout. We find that composition backfires on every resource axis we measure. Compression costs a recursive reader 4.1 accuracy points (95\% CI [$-$7.2, $-$1.1], p = 0.008) while raising its input-token load by 20--56\% on five of seven models and its latency on all seven: a reader working from a lossy document re-reads and re-queries, more than doubling contact with the 30-call tool budget (59 $\rightarrow$ 121 runs, p = 9$\times$10$^{-9}$). That mediation is tested rather than assumed; it accounts for about half the penalty. Recursion's own large pooled advantage (+17.3 points) is window relief, not comprehension. We stratify on the API's recorded overflow outcome, then audit that outcome against cross-model token anchors: the same prompt's recorded token count on models whose window it fit. The audit exposes 39 mis-recorded outcomes (0.7\% of runs), almost all at the stratum boundary. On documents that overflow, recursion answers 77.7\%, +52.7 points over the 25\% chance floor; on documents that genuinely fit, it is equivalent to a direct call within $\pm$5 points (TOST p = 0.003; point estimate $-$0.4) at 1.7--15.7$\times$ the latency, and an apparent length gradient in the un-audited data collapses once the silent overflows are reclassified. Compression on a direct call rescues 38\% of over-window documents but scores 27.6\% there, indistinguishable from guessing, while costing 4.5 points on documents that already fit. Pooled comparisons of long-context methods are therefore length-distribution mixtures, not capability measurements.
Chat is not available.
Successful Page Load