One Coordinate Away from the Right Answer
Abstract
Small language models often compute whether a conclusion follows from its premises, then answer as if they had not. A linear probe reads deductive validity from their mid-depth activations far more reliably than the models’ own answers separate the same items. We intervene twice: deleting the one premise a proof requires, and amplifying each item’s own coordinate along the model’s validity direction inside the forward pass. The internal verdict tracks the deleted evidence while the answer does not, and the amplification reconnects them. Which models show the dissociation, and which are repairable, is family-dependent: by 7–8B the dissociation is gone in every family we test, while the gap between what the probe reads and what the answers use persists, and the amplification still closes much of it. We release balanced splits, per-item scores, and code.