The Butterfly Effect in Reasoning: Branch-Structured Distributional Memory for Stochastic Agents
Abstract
Memory-augmented reasoning agents typically store a single realized reasoning trajectory as the unit of experience. This abstraction is brittle under stochastic decoding: the same task can induce multiple plausible reasoning paths with different decompositions, assumptions, and failure modes. We introduce Branch-Structured Distributional Memory (BranchDIME), a framework for constructing memory from branch-structured trajectory sets rather than isolated traces. BranchDIME samples short reasoning prefixes, clusters them into prefix-induced branches, expands representative prefixes into full trajectories, and distills the resulting set into reusable memories containing positive strategies, anti-patterns, and contrastive rules. Across GPQA-DIAMOND and MATH500, BranchDIME improves over single-trajectory memory and alternative trajectory-selection baselines at comparable inference latency, yielding relative gains of 8.7% and 2.2% over single-trajectory memory, respectively. Coverage analysis shows that BranchDIME spans all discovered branches, compared with 20.0% coverage for consensus selection and 64.9% for random selection, while also achieving the highest downstream accuracy. BranchDIME further transfers across tasks and models, improving over single-trajectory memory by up to 15.4% on AIME25 and OlymMATH. These results support a shift from trajectory as experience to trajectory distribution as experience: memory should capture the local reasoning landscape induced by stochastic generation, not merely one sampled path.