ReCon: Toward Balanced Learning under Inter-Context and Context-Memory Conflicts
Abstract
In retrieval-augmented generation (RAG), knowledge conflicts arise when retrieved contexts disagree with each other or with the model's parametric memory. Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a standard approach to address such conflicts. However, existing RLVR-based methods are typically formulated under a single-conflict assumption. As a result, the type with stronger advantage signals dominates policy updates, leaving the other under-optimized and thus yielding imbalanced performance across conflict types. To address this challenge, we propose ReCon, a type-aware RLVR framework for balanced learning across Inter-Context and Context-Memory conflicts. Specifically, ReCon computes a step-level imbalance signal from type-wise cumulative advantages and triggers resampling on the weaker conflict type. Instead of resampling from scratch, ReCon branches from informative failures. It selects an informative failed trajectory using uncertainty and closeness signals, then branches at an entropy-jump point, concentrating additional rollouts on critical reasoning decisions. Experiments show that ReCon yields stronger and more balanced performance across knowledge-conflict and multi-hop QA benchmarks, and also generalizing well under mixed-conflict settings.