RealDev-QA: Trajectory-Level Diagnosis for Developer RAG Under Real-World Noises
Abstract
Real-world developer help-seeking is rarely a clean search query: requests mix stack traces, failed attempts, environment constraints, and mistaken hypotheses, while relevant-looking documents may be invalid under the user's specific runtime settings. We introduce RealDev-QA, a trajectory-supervised benchmark for developer RAG under diagnostic query noise and evidence conflict. RealDev-QA contains 1,233 post-June-2024 multi-hop questions derived from public developer discussions, each paired with decomposed sub-queries, dependency edges, per-step gold evidence, verified reasoning graphs, and adversarial hard negatives drawn from real documentation, release notes, and verified issues. Its construction uses answer-conditioned evidence tracing to ensure each retained instance has a verifiable evidence path; evaluation then removes the resolution and tests whether systems can recover and use that path from the diagnostic query alone. Across static, graph-based, and agentic RAG systems, comprehensive evaluations show that retrieval helps but does not solve RealDev-QA: the best human-audited accuracy reaches only 24.7\%. Trajectory diagnostics show that failures are typically decided early: over 78\% of incorrect trajectories miss gold evidence in the first retrieval step, and first-query precision predicts downstream correctness beyond aggregate evidence coverage—broad initial queries admit hard negatives that persist and pollute later evidence navigation. Low correctness decomposes into five pipeline-level failure modes, while retrieval, automatic, and human-grounded metrics yield divergent system rankings largely masked by aggregate evaluations on standard benchmarks. RealDev-QA reframes developer RAG as evidence-path control: deciding where to start, what to connect, what to reject, and how to recover.