Open Problems in LLM Self-Prediction
Abstract
We tested how well LLMs can predict what they will do in long-horizon agentic environments. We found LLMs are poorly calibrated and overconfident about how well they will perform; stronger models rise to their own inflated expectations, which makes them merely appear to be better self-forecasters. We release a detailed evaluation suite ([code and data link redacted for double-blind review; see supplementary]), so that self-predictive ability can be tracked in future models: a sudden step in self-knowledge above our baseline could be a warning sign for long-horizon scheming, since self-knowledge plausibly is a prerequisite for it. Models are closer to unbiased on stylistic dimensions than on capability, though no better at discriminating between tasks, and when asked to predict misaligned behavior they report themselves as safer than an anonymous peer.