Linking Trajectory Drift to Representational Degradation in Sequential LLM Forecasting
Abstract
As LLM agents are deployed in open-ended, long-horizon environments, sequential self-feedback introduces critical reliability and oversight risks that standard endpoint evaluations fail to detect. We investigate this vulnerability by holding the forecasting task and external evidence stream fixed, systematically varying whether models observe no prior forecasts, their immediately preceding forecast, their full forecast history, or full history with model-generated rationales. Across 100 Kalshi markets with 30 timestamped news steps, we evaluate Qwen2.5-7B-Instruct, Llama-3.1-8B-Instruct, and Gemma3-4B-IT. Exposing a model to prior forecast(s) increases drift across all three models, but the resulting performance failures differ sharply: Qwen's multi-turn bare Brier score worsens from 0.188 to 0.234 while AUC rises from 0.790 to 0.850, whereas Llama loses both calibration and discrimination, with AUC falling from 0.899 to 0.806. Linear probes show that self-generated history reduces outcome decodability from residual-stream activations across all three architectures. Counterfactual rewrites further show that prior forecasts causally anchor subsequent predictions, with Qwen exhibiting stronger anchoring to inflated than deflated forecasts (upward drift). Rationales do not consistently mitigate these effects, partially restoring decodability in Qwen while further degrading it in Llama. Our results show that accumulated self-generated state can systematically distort sequential belief updating in persistent agents and that these failures are not captured by endpoint accuracy alone.