Long Documents, Fixed Budgets: Context and Turn Economics of Post-Training a Document-Reasoning Agent
Abstract
Professional documents are a long-context workload that retrieval-style benchmarks do not capture: a task's evidence is spread across a source document of median length 27 pages (mean 48.8, maximum 873), a third of tasks rest mainly on charts, drawings, and maps, and 70.9% require two or more reasoning hops. We study what happens to a foundation model's use of context when it is post-trained on such documents under a fixed 65,536-token window, and what happens to its use of turns when the resulting policy is run as a long-horizon agent. The training corpus is 1,465 expert-authored tasks with atomic rubrics drawn from a ten-domain enterprise collection; the recipe is privileged-context on-policy self-distillation followed by criterion-reward reinforcement learning, both with rank-32 adapters on Kimi-K2.7. Training on the raw PDFs would have required a context window of more than 256K tokens; we normalize every document to extracted text plus selected page images, reserve the response budget against the longer privileged prompt before generation, and reject truncated samples rather than train on them. On the 100-task GDP.pdf holdout, strict pass rises from 11.0% to 24.5% while mean tokens generated stay flat (9,803 to 9,733), so the gain is accuracy per token, not length; GPT-5.6 Sol's leaderboard strict pass is 30.7% at a measured 3,422 tokens, which shows how much of the efficiency gap remains. On 220 out-of-distribution GDPval tasks, messages per task fall from 70.9 at base to 58.2 after the third round while mean rubric reward rises from 0.571 to 0.746 and zero-credit tasks fall from 54 to 10; paired trajectories show a context-thrashing retry loop (roughly 110 messages, the same statute re-extracted sixteen times) replaced by a single up-front census of sources. We argue that long-context evaluation of agents should report accuracy jointly with tokens and turns, and that datasets of long documents should ship text and image renderings alongside the PDF.