Safety Refusal Is Upstream of Knowledge Abstention
Abstract
Some work studies safety refusal, other work studies abstention on entities a model does not know, and which of them governs when both apply to the same prompt has not been asked. We built a corpus where they compete. On the surface safety wins: where both grounds hold, the model gives the safety reason 76.7% of the time and the epistemic one 22.2%. Ablating each direction shows the two are not competing on equal terms. Removing the safety direction does not let abstention take over. Withholding collapses instead, and the model starts inventing details about an entity it has never seen, in 69% of answers against 9% before. Removing the abstention direction leaves safety refusal where it was, indistinguishable from a norm-matched random direction. The dependence runs one way.