Stopping Superseded Tool Calls Without Dropping Unrelated Work
Abstract
A user correction can invalidate an in-flight tool call while leaving unrelated work valid. We separate execution freshness from instruction understanding in two controlled studies. First, frozen model proposals are replayed under five execution rules and four schedules over 80 histories. For Qwen2.5-1.5B, task-scoped atomic validation removes 62 stale in-flight commits and preserves 128 useful commits, versus 65 under global validation, but still accepts 23 wrong-value commits. A smaller model fails the strict JSON interface, making its zero-commit result vacuous. Second, 1,134 fresh generations across Qwen2.5-1.5B and Qwen3-1.7B test whether extracting chronological edits for a reversible executor improves revision understanding. On 192 histories, event extraction scores 49 and 112 exact states, versus 116 and 95 for an explicitly instructed direct-state control. The evidence supports separating commit validity, interface validity and semantic accuracy when evaluating interruption handling. These are synthetic diagnostic studies using established concurrency principles, not a new transaction algorithm or deployed conversational benchmark.