Evaluating Next-Frame Predictions, Long-Term Predictions, and Protocol Timing in Operando Microscopy
Abstract
Experimental movies often change slowly, so copying the latest frame is already a strong next-frame baseline. We examine this problem using operando interferometric scattering microscopy of four graphite particles measured at 25 and 45\,°C. Across 24 model configurations, 15 outperform persistence for one frame: 10 of 12 residual-target and five of 12 direct-target models. The best mean absolute error (MAE) ratio is 0.811, compared with 0.880 for a fitted five-parameter linear autoregressor. The deep model improves on the linear baseline, but simple local dynamics explain much of the next-frame change. The same pattern appears when each particle is left out of training in turn. Long predictions give a different result. No model average outperforms persistence after 32, 128, 256, or 512 consecutive predictions. Repeating the best long-rollout configuration with four initialization seeds gives no seed average or particle--seed result below persistence at any reported horizon. At 512 steps, none of 96 particle-level comparisons outperforms persistence, and five predictions become non-finite. A ConvLSTM uses voltage and current, as shown by zero and shuffle tests, but delaying these inputs by 16 to 512 frames changes its one-step ratio by only 0.0005 to 0.0126. At the largest delay, voltage and current differ from the measured inputs by 0.83 and 0.62 standard deviations on average. We also show that updating the persistence reference or recalibrating optical thresholds during prediction can hide error. Long predictions should therefore use one fixed baseline and should test both the values and timing of external inputs.