Grounding Decay in Long-Sequence Chart Extraction: Position-Dependent Error in Vision-Language Models
Abstract
Vision-language models (VLMs) are increasingly deployed in settings that require long, structured outputs grounded in a single image, yet little is known about how reliably visual grounding survives extended autoregressive decoding. We study this question in scientific chart digitization, a task where a model must transcribe hundreds of numerical coordinates from one figure into a deeply nested schema, and where every output token has a verifiable ground-truth referent in the image. We build a two-stage cascaded pipeline that isolates individual graph panels to maximize Vision Transformer (ViT) patch allocation and then extracts axes, legends, data points, and error bounds into a strict machine-readable YAML schema using a fine-tuned Gemma-4-31B. The pipeline is competitive with proprietary zero-shot baselines on chart types with short output horizons, reaching a Normalized Mean Absolute Error (NMAE) of 1.55% on bar charts and 7.97% on line charts, but degrades sharply on dense scatter plots that require thousands of output tokens. Partitioning each scatter plot into four X-axis quartiles, which approximate early-to-late generation because the model emits coordinates in ascending X order, mean Jensen-Shannon Divergence rises from 0.571 in the first quartile to 0.847 in the fourth. Against a subsampling null that controls for the sparsity of late slices, divergence above chance rises by 122% across the same range. We characterize this as position-dependent grounding decay, in the descriptive sense that extracted values diverge further from the visual evidence the later they are emitted, without any accompanying failure of schema compliance. Because ascending-X decoding ties X position to generation index, we separate them with a within-panel control on line charts, where later-emitted series span the same X range but occupy later generation indices. Later series are less accurate in 69 of 85 panels, with a Wilcoxon signed-rank p-value below 0.00000001. Output position is therefore an independent predictor of extraction error, and we specify in-distribution experiments designed to isolate the underlying mechanism.