Context-Answer Metamodels: Forecasting LLM Activations Before the First Answer Token
Abstract
Under stochastic decoding, a single prompt can elicit many different responses from a language model. Characterizing this variation typically requires many expensive autoregressive rollouts, each revealing only one realization. We ask whether the activations produced while reading the prompt can instead forecast the model's later internal states. We train context--answer metamodels that map a prompt activation sequence to point or distribution forecasts of mean-pooled answer activations. We find that (i) the four metamodel types forecast mean-pooled answer activations and their rollout variability with high fidelity; (ii) performance scales steadily with prompt count through 498,600 training prompts and generalizes zero-shot to out-of-domain tasks; (iii) these metamodels retain behavior-predictive signal; and (iv) prompt activations from a weaker model predict mean-pooled answer activations in a stronger model, suggesting cross-model transfer from weaker to stronger models. Together, these results suggest that prompt activation traces can support scalable forecasts of behaviors in LLM answers before the first answer token is generated.