Belief Context Graph: A Revisable State for Long-Horizon Agents
Abstract
Long-horizon language-model agents need more than loops that repeat work and workflow graphs that organize dependencies: they need an explicit, revisable state of what remains supported as observations accumulate and conflict. Raw trajectories mix external evidence, model hypotheses, failed actions, corrections, and decisions. Replaying this history increases cost, while truncation, retrieval, and flat summarization can erase provenance or uncertainty. We introduce Belief Context Graph (BCG), a context-management method that transforms the model context into an evidence-grounded graph of beliefs and decisions. Nodes preserve source role, stance, temporal metadata, evidence provenance, merge lineage, and inspectable confidence signals; typed edges express dependency, supplementation, and contradiction. With GPT-5.6-Luna on BrowseComp and BrowseComp-ZH, changing only context management from Default to BCG improves accuracy from 33.33% to 37.12% and from 49.48% to 59.17%, while reducing mean total token use by 16.2% and 9.5%, including graph-construction cost. Results across the two evaluated agent models show that explicit belief maintenance can improve the accuracy--resource trade-off for long-horizon agents.