SWE-Touch: Benchmarking Coding Agents When Users Touch the Code
Abstract
Real-world software development requires coding agents to share workspaces with users who may modify code during a task, yet repository-level benchmarks generally evaluate agents working alone or restrict users to messages. We introduce SWE-Touch, a framework that stress-tests shared-workspace collaboration with validated Counter-Edits: plausible changes to task-relevant code that conflict with task completion. SWE-Touch mines task-critical regions from multiple repair trajectories, uses a separate User Patch Generator to construct small, locally plausible edits, validates that they conflict with the task, and injects them with contextual messages when agents reach the relevant code. Across 9 models, Counter-Edit lowers average resolve rate by 7.7 percentage points on SWE-bench Verified, with degradation persisting on the longer-horizon SWE-Bench Pro and DeepSWE tasks. Trajectory analysis shows that agents may retain conflicting code or replace it without sufficiently re-inspecting the repository and validating the revised behavior. These findings show that strong autonomous performance does not ensure the state awareness needed for shared-workspace collaboration, motivating better change detection, conflict reconciliation, and targeted validation.