When Replay Agreement Misleads: Capture Loss and Cancellation in Local Agents
Abstract
Local tool-using language models can show almost identical aggregate decision and path agreement while changing execution on individual cases. We analyze 800 retained episodes of a Qwen 2.5 7B configuration over 100 synthetic financial cases, with eight repeats per case. Both agreement measures equal 99.625%, yet one of 97 unanimous-decision groups changes its tool path. An opposite-sign case exactly cancels this case's contribution to the aggregate gap. This is a counterexample to interpreting a zero gap as an absence of hidden path variation, not evidence that the variation is harmful. A separate 48-episode diagnostic study of manifest-checked Qwen 3.5 and Granite 4.1 configurations achieves perfect decision and strong-path agreement in all 16 groups. Consistently deleting recorded calls preserves these scores while corrupting the evidence. A dispatcher-side comparison detects the injected omissions; required-tool-name checks alone miss omissions of optional calls. We provide synthetic traces and an offline audit. These are bounded counterexamples about measurement, not financial-accuracy, efficiency, or model-ranking results.