Can Recurrent Models Recover Ground-Truth Temporal Lags?
Abstract
Neural encoding models test whether artificial neural network representations can predict biological responses to the same stimuli. Strong prediction shows that a representation contains response-relevant information, but the lag with the highest encoding score may not recover the planted delay that generated the response. We test this distinction with DelayBench, a controlled language benchmark in which both the response-generating features and their token-scale delays are known. A gated recurrent unit (GRU) learns next-token prediction, and ridge regression maps its hidden states to held-out synthetic responses at candidate delays. Across 12 seeds, the trained GRU nearly matches the oracle representation that generates the responses. Its mean peak Pearson correlation is 0.970±0.001, compared with 0.972±0.001 for the oracle. Training improves prediction: the trained GRU exceeds an untrained GRU by 0.065±0.005 and its context-free token embeddings by 0.136±0.002. However, the trained GRU recovers the exact planted delay for only 54.2%±8.2% of channel-seed pairs, with a mean absolute error of 0.71±0.17 tokens. Control representations also remain predictive: an untrained GRU reaches r=0.906, lexical identity reaches r=0.836, and position reaches r=0.645, while a response-shuffled null is near zero at r=0.028. Recurrent states retain information across successive tokens, allowing several candidate lags to predict the same response accurately. Encoding performance and temporal identifiability are therefore distinct properties. \bench provides a ground-truth test for lagged ANN-to-brain analyses before researchers assign temporal interpretations to human recordings.