On Length Bias in EEG-to-Text Decoding
Abstract
Sentence-level EEG-to-text decoding models typically pad variable-duration neural signals to a fixed-length window, creating a structural confound where segment duration correlates with sentence identity independently of neural content, a problem we term length bias. Despite being acknowledged in prior work, the prevalence, magnitude, and architectural dependence of this effect have never been systematically quantified. We introduce length-matched noise baselines that isolate this confound by replacing only the true signal region with Gaussian noise while preserving zero-padding, and apply them across three English listening EEG datasets under a strict subject-held-out protocol. Evaluating ten sentence classifiers and three generation models based on the current state-of-the-art, we show that reported decoding scores are substantially inflated by length cues across datasets, with the degree of inflation varying markedly by architecture. Among classifiers, BIOT alone is unaffected by length-matched noise. A systematic decomposition of its components identifies Short-Fourier-Transform (STFT) tokenisation as the bigger source of this invariance. We transfer this finding into two new generation variants that largely eliminate encoder-side length bias, enabling a controlled analysis of where residual bias originates. A classification-to-generation evaluation further reveals that leading generation models operate largely as sentence retrieval systems, with a linear classifier recovering nearly all of their reported BLEU-4 scores. Together, these results establish that length bias is a pervasive and previously unquantified failure mode in EEG-to-text evaluation, and provide both diagnostic tools and architectural directions for future work.