Predicting Response Semantics Before Generation: Identification and Causal Intervention
Abstract
Previous work has shown that the semantic content of future responses can be predicted from model states before generation, but prediction performance itself establishes only that a predictive relationship exists between the model state and future semantic content. We more rigorously evaluate this prediction using re- sponse level identification and multiple controls. Using Llama-3.1-8B-Instruct and CommonGen, we sampled ten responses for each prompt and averaged their response embeddings to form a response centroid as the operational target for response semantics. A linear mapping from the hidden state at the last prompt token identified the corresponding response centroid among 32,620 candidates with 84.8% Recall@1. A linear mapping from SimCSE prompt embeddings to response centroids was also predictive but achieved 67.0% Recall@1, while direct prompt response similarity achieved 1.4%. Identification remained high when competing candidates were associated with other prompts containing all concepts specified in the source prompt, and after removing the first response token. Together, these controls show that the high identification performance extends beyond the immedi- ately upcoming first token, persists beyond simple matching of concepts explicitly specified in the prompt, and is not fully reproduced by a linear mapping from SimCSE prompt embeddings to response centroids. In addition, moving the same hidden state toward donor states produced a small but direction specific semantic shift relative to a norm matched random perturbation. These results show that, under the conditions studied here, subsequent response semantics can be predicted and identified with high accuracy from the hidden state before generation, and that a specific localized intervention on the same state causally influences subsequent response semantics.