Visual Anchoring for Scenario-Guided Forecasting
Abstract
A multimodal large language model (MLLM) forecaster takes a forward-looking scenario as a soft constraint that decays with horizon. The decay is expected, but its rate is uncharacterized and untunable at inference time. A practitioner running a twelve-month stress test cannot tell whether the scenario binds for twelve months or twelve days. Existing accounts of autoregressive error, such as exposure bias and hallucination snowballing, bound a single rollout against ground truth and offer no inference-time lever. Under three assumptions, we prove a closed-form Lipschitz upper envelope on the per-step scenario drift between two paired-seed forecasts differing only in scenario text. Two identifiable parameters partition behavior into bounded, linear, and exponential regimes. A sister within-chunk bound holds under any chunk partition. It motivates a stabilization heuristic that we test and falsify: periodic re-injection of the scenario enlarges drift relative to a single injection, and length-matched random text shrinks it. The envelope is respected; the heuristic is not. A pre-registered attention diagnostic across the open-weight panel returns architecturally heterogeneous verdicts that the recurrence framing covers uniformly while no single attention-dilution mechanism does. In place of text re-injection, we propose Visan (Visual Anchoring), a chart-anchored multimodal prompt that renders the lookback, the scenario, and the historical context as a single image. Across a panel of MLLMs and two long-horizon scenario-conditioned forecasting benchmarks, Visan reduces forecast error on the majority of MLLMs.