Identifying Trajectory Specific Command Uptake in Software Engineering Agents: Evidence from Cue Exclusion Controls
Abstract
Software agent memory is often evaluated by whether the task succeeds. That outcome leaves open whether memory changed the agent's actions. We introduce a provenance based test for uptake from a repair trajectory. For each pair of prior and target issues, we isolate normalized commands that appear in the trajectory and in neither issue statement. NONE and ISSUE omit these commands from the prompt, whereas TRACE, ACTION, and DIGEST expose them. Across 95 SWE-ContextBench targets and three OpenHands configurations, any command reuse is 10.88 to 17.54 percentage points higher under the trajectory renderings than under NONE. On the two configurations with ISSUE, reuse is 1.05 points above NONE under ISSUE and 12.11 to 20.00 points above ISSUE under the trajectory renderings. For any command reuse, all three renderer comparisons with NONE and all three with ISSUE survive their respective Holm corrections. The excess exact matches provide an auditable behavioral signature of trajectory uptake beyond the prior issue. Complete episodes also have full resolution rates 5.26 to 9.12 points above NONE. On the configurations with ISSUE, point estimates place ISSUE between NONE and the renderings, but all four adjusted outcome tests have (p>.30). The command evidence identifies uptake more clearly than the outcome evidence separates the prior issue from the trajectory.