Overthinking as a Symptom of Knowledge Conflict: Understanding and Detecting LLM's Hallucinations in Retrieval-Augmented Question Answering
Abstract
Retrieval-augmented generation (RAG) is increasingly used to ground Large language Models (LLMs) with up-to-date or private information, and has become a key component for knowledge-intensive LLM systems (e.g., LLM agents). However, RAG with LLMs often hallucinate due to unreliable retrieved information or outdated internal knowledge, making hallucination detection crucial for reliable LLM workflows. Recently, reasoning-intensive LLMs that produce explicit reasoning trajectories before answering have shown promise in RAG by examining and synthesizing retrieved evidence. Yet, they often suffer from overthinking: generating redundant or unproductive reasoning with limited performance gains. While prior work mainly studies overthinking in closed-book reasoning tasks that mainly rely on internal knowledge, overthinking in RAG-based question answering (QA) that requires open-book reasoning over both internal knowledge and external context remains underexplored. Thereby, we conduct empirical studies to investigate the relationship between overthinking and hallucination in RAG-based QA and find: (1) Overthinking in reasoning LLMs produces lengthy reasoning trajectories that are correlated with hallucinations; (2) This overthinking is often associated with knowledge conflict, where LLMs struggle to reconcile retrieved context with internal knowledge. Based on our findings, we propose a knowledge conflict-aware hallucination detection framework. We train a classifier to identify the three knowledge conflict reasoning patterns, design a knowledge conflict score to measure knowledge conflict severity, and apply tailored detection strategies with different prompts for different conflict levels. Unlike existing methods, our approach requires neither LLM parameter access nor retrieval reliability, and only needs the LLM output in a single run. Experiments on three QA datasets and three LLMs confirm our framework's effectiveness.