Pdb Or Not To Be: A Budget-Matched Comparison of Execution Interfaces for Software Agents
Abstract
Interactive debuggers are indispensable to human software development, yet coding agents overwhelmingly debug through unstructured shell commands, using common bash tools, test scripts, and print statements, and re-running code rather than stepping through a live debugger. On SWE-Bench Verified, we compare a debugger-only agent against a bash-only agent to assess whether an effective bug-fixing agent can be built without raw shell access. On overall pass rate, we find that the shell agent is actually the stronger default across both models we study, though the debugger's relative value is highly model-dependent. Additionally, we isolate each tool's contribution to the diversity of their union. The notion that diverse scaffolds can solve complementary tasks is well-explored; however, we show that, while the bash and debugger interfaces solve partially disjoint instances in a pass@1 setting, the debugger in isolation adds less diversity to a shell-agent portfolio than an additional shell-agent run does. We establish this with a budget-matched cross-scaffold best@2 that disentangles the genuinely unique contributions of each architecture from mere stochastic variance in task coverage across runs.