Mitigating Knowledge Conflicts in Retrieval-Augmented Generation via Inference-Time Representation Editing
Abstract
Retrieval-Augmented Generation (RAG) has been widely adopted to enable Large Language Models (LLMs) to ground their responses in external knowledge sources. Nonetheless, recent studies show that conflicts between the retrieved external knowledge and the model’s parametric knowledge can lead to hallucinatory outputs, and this problem is exacerbated when the retrieved documents contain noise. In this work, we propose Conflict-Aware Representation Editing (CARE), an inference-time representation editing method for improving LLM robustness under noisy knowledge conflict settings. CARE learns a latent editing direction from intermediate representations associated with correct and incorrect generations, and uses this direction to mitigate conflict-associated activation patterns during inference. We evaluate CARE across Question Answering (QA) benchmarks and LLMs, showing improvements over retrieval-based and conflict-mitigation baselines, especially under noisy retrieval conditions.