Which Predictions Are Worth Updating? Budgeted Predictive Knowledge for Long-Horizon Agents
Abstract
Long-horizon agents rely on predictions about their environment, but these predictions can become stale as the environment changes. When fresh evidence is limited, which predictions should an agent update? We study this question through budgeted predictive-knowledge updates: selecting a fixed number of stored predictions to revise before a fixed LLM policy continues acting. We develop three selection strategies that prioritize predictions by change risk and downstream importance, adapt priorities using execution feedback and unfinished work, or select updates jointly to account for shared prerequisites. Across 1,872 episodes in two controlled task families and four LLM policies, combining change risk with downstream importance improves the fraction of jobs completed over random selection by 7.6–10.9 percentage points on the primary model. A follow-up study shows that accounting for remaining dependencies improves progress under execution feedback, without establishing an additional benefit from joint selection. These findings show how dependency information can guide limited knowledge updates, while revealing a persistent gap between better intermediate progress and reliable task completion.