SWE-Git-Bench: A Focused Worktree-Level Benchmark for Real Merge Conflict Resolution
Wei Zhang ⋅ Jian Yang ⋅ jiajun wu ⋅ Zidan Tang ⋅ Siwei Wu
Abstract
Modern coding agents operate inside live \texttt{git} worktrees: they edit files, pull upstream changes, and must leave the repository in a mergeable state. Yet current code benchmarks focus on issue resolution or synthesis from specification, not on the merge-resolution primitive an agent must invoke after \texttt{git} emits conflict markers. We introduce \textbf{SWE-Git-Bench}, a focused high-fidelity benchmark of 135 manually audited real conflicts spanning 185 files from 23 repositories. Across 36 contemporary LLMs, exact reproduction of the maintainer-committed file peaks at only \TopWFEM{}\%, while top systems still cluster around $86$--$90\%$ edit similarity, exposing a near-repair regime where outputs look close but fail to reproduce the committed resolution. A human-checked audit of 80 whole-file non-EM outputs with ES $\ge 0.99$ finds that 50.0\% are still likely unsafe, so similarity cannot be treated as a semantic pass rate. Conflict-block prompting improves edit similarity by \AvgDeltaES{}\,pp on average (median $+16.1$\,pp) while leaving exact match effectively unchanged. We release the dataset, harness, full prediction matrix, difficulty calibration, and exploratory failure analysis to support reproducible evaluation of this worktree-replayed, file-scored resolver primitive.
Chat is not available.
Successful Page Load