The Override Gap: A Magnitude Account of Knowledge Conflict Failure in Hypernetwork-Based Instant LLM Adaptation
Abstract
Hypernetwork-based methods such as Doc-to-LoRA internalize a document into an LLM's weights in a single forward pass, but they fail systematically when the document contradicts pretraining knowledge: accuracy drops to 46.4% on the deepest conflicts. We show that this failure is primarily a magnitude problem rather than a representational one. The generated adapter reaches the relevant layers, but its adapter margin is not calibrated to the strength of the contradicted prior; as the pretrained margin grows, deep conflicts lose the override competition. This account predicts that failure should track prior strength. Sorting 194 conflicts by the base model's log-probability on the contradicted fact, baseline accuracy falls from 68% on weak-prior questions to 16% on strong-prior questions, a 52 percentage-point gap. We propose two training-free corrections. Selective Layer Boosting (SLB) scales the adapter at its highest-activity layers, and Conflict-Aware Internalization (CA) applies stronger boosting only when the base model appears confident. Together they raise deep-conflict accuracy from 46.4% to 71.0% on Gemma-2B and from 53.6% to 72.5% on Mistral-7B, while preserving novel-knowledge recall at 97.1%. They outperform vanilla retrieval-augmented generation on Gemma medium conflicts by 18 percentage points, though explicit conflict-aware prompts remain stronger when the user already knows a conflict exists. We release KID-Bench, a 489-question benchmark that separates novel recall, cross-knowledge combination, and prior-graded conflicts.