Do Factual Edits in LLM Embeddings Propagate Across Two Hops?
Abstract
Whether an LLM's embedding space connects related facts, or stores each fact independently, remains an open question. To investigate this, we design a controlled embedding-level intervention. We freeze the model and learn entity-specific residuals for the individual facts underlying two-hop examples from MQuAKE-CF dataset using Llama-3.1-8B and Qwen2.5-7B. First, we test whether facts that are successfully edited individually also carry through when they are combined. The models can answer many of the edited single-hop facts, but rarely produce the corresponding answer when the two facts must be combined: direct two-hop accuracy is only 1.1\% and 1.4\%. Second, we test whether this failure is caused in part by the intermediate entity's edit not being available during the two-hop computation. We let the model generate the intermediate entity and then explicitly use it in the second-hop prompt, allowing its residual to be activated. For the same generated entity and second-hop prompt, activating the residual increases accuracy from 0.6\% to 15.5\% and from 1.3\% to 9.2\%. However, access alone is not sufficient: even when the correct intermediate entity is generated and its residual is active, final accuracy remains very low. Finally, to test whether this failure is specific to embedding-level editing, we introduce the same facts using LoRA updates to the transformer layers, but two-hop performance improves only modestly. Overall, successful individual embedding edits do not imply that the embedding representation supports the corresponding two-hop factual relation.