Study or Search? Designing Efficient and Effective Coding Agents for Downstream Tasks
Abstract
AI coding agents repeatedly reconstruct knowledge of unfamiliar repositories through test-time search, but it remains unclear whether studying a codebase before a given downstream task produces reusable knowledge that improves later answers. This question is difficult to evaluate because a plausible preparation artifact does not show that the agent actually read its corpus, and additional search or reasoning can mask a weak preparation method. We investigate these issues with a controlled evaluation of three pre-task conditions: No guide (base), a one-pass frozen grounded cheatsheet, and a six stage active-study procedure. We cross these conditions with three repository-search budgets, two thinking modes, and 30 questions about DSPy from StudyBench, a benchmark for evaluating how well AI agents can answer expert-level coding questions about real open-source codebases. The resulting experiment contains 540 planned trajectories, of which 525 complete. Preparation is isolated from benchmark questions and grading material, source interaction is validated before each guide is frozen, and a separate claim-level grader is blinded to treatment and private reasoning. Active Study has the highest aggregate lenient score, the simpler cheatsheet leads every thinking-off setting, and no guide (base) catches or exceeds both prepared conditions when search and thinking are combined. These findings show that preparation quality, inference resources, and reliability must be evaluated jointly, and that a more elaborate study process does not necessarily yield a more useful preparation artifact or better downstream answers.