When Search Succeeds but Agents Fail: Diagnosing and Mitigating Retrieval Myopia
Hyeondong Woo ⋅ Jongwon Ryu ⋅ Jinkwon Hwang ⋅ Junyeong Kim
Abstract
RL-trained search agents interleave retrieval with reasoning to answer knowledge-intensive questions, achieving strong end-to-end accuracy on multi-hop QA benchmarks. Yet end-to-end accuracy conflates two distinct abilities: searching for relevant documents and reading what they contain. Existing evaluations do not measure whether the agent reads the documents it retrieves. We design a controlled evaluation that force-injects gold-bearing documents into each agent's retrieval pipeline, isolating reading from retrieval. Across four RL search agents on three multi-hop QA benchmarks, up to $46.2\%$ of samples are answered incorrectly even when the gold answer is literally present in the retrieved documents: retrieval succeeds but reading fails. We term this phenomenon Retrieval Myopia (RM). Re-presenting the same naturally retrieved documents to the same model in a plain QA format, without the accumulated interaction history or any weight updates, recovers up to $48.9\%$ of RM cases. Guided by this diagnosis, we propose CWFT (Confidence-Weighted Format Transfer), a training-free pipeline that anchors the agent's prediction against $k$ confidence-weighted clean-format re-reads. CWFT consistently improves accuracy across all evaluated system-dataset configurations on three multi-hop QA benchmarks, while retaining approximately $99\%$ of initially correct answers.
Chat is not available.
Successful Page Load