Are Correct Causal Answers Decodable Even When Model Behavior Is Weak?
Abstract
Does a language model contain a representation of the correct answer to a causal question even when it does not give that answer? We test this question on the Corr2Cause benchmark using two frozen instruction-tuned models. From each model's final layer, we extract one hidden state after only the premise and one after the hypothesis. A linear probe is trained on 4,000 balanced training examples to predict the label. On the test set, hypothesis-end balanced accuracy is 81.3\% for Llama and 80.6\% for Qwen. Both results exceed the performance of the models' behavioral answers. Llama itself predicts No for every test item (50.0\% balanced accuracy), providing a dissociation between the behavior and the representation. This is also seen with Qwen, as its answers are near chance (52.9\%, 95\% interval 50.5-55.3\%). These findings show that the answer is available to a linear decoder in both models, even when it is not reflected in their answers.