Actions Without Reasons: Filling the Reasoning Gap in a Large-Scale Public Dataset of Real Coding-Agent Trajectories
Abstract
Deployed coding agents leave behind real trajectories that record actions without reasons. Across 26,999 sessions and 1.65M assistant turns we measure 15.3 non-empty reasoning blocks per 100 turns, so over 84% of assistant turns carry no reasoning supervision. Training on such data teaches imitation of what was done, not inference of why. We propose action-grounded reasoning synthesis: a writer sees the context and the kind of action taken but never its arguments. Of several candidates per position we keep the one that most raises a frozen student's log-probability of the true action, ranked against each other rather than against a fixed threshold. The method needs no existing thinking blocks as input, so it applies to any trajectory, including sessions that carry none at all. Filtering by whether a rationale raises the probability of the true action favors hindsight: rationales written with the action in view pass such a filter at 49.6% against 1.6% for action-blind ones, so any filter that rewards "leads to the correct answer" ends up selecting for hindsight. Under matched-budget training, reasoning supervision lowers held-out real-thinking perplexity by 18.4%, and filling empty positions improves held-out action log-probability by 3.9% on average across seed pairs, above seed variation. Synthesis adds decision signal rather than reproducing the text distribution of real thinking. We release Trajector-2.5B: 26,999 consented sessions (2.50B tokens), redacted, decontaminated, and duplicate-labelled.