Before Cross-Site Transfer: Reliability of Tree and Recurrent Models with Limited Wastewater Surveillance Data
Abstract
Problem and contribution. Transferring forecasting models to observation-sparse domains requires a credible local reference at the target. For irregular, non-stationary series, that reference is fragile: it varies with evaluation discipline, forecast origin, and which observations survive, not merely how many. Wastewater viral concentrations are indirect population-level signals shaped by shedding, transport, dilution, and assay, so communities within the same surveillance network can provide models with materially different forecast-time information. Because direct rich–poor comparison confounds sparsity with catchment, epidemic, and measurement differences, we degrade a densely observed community under controlled conditions while holding the forecast task fixed. Evaluation design. We predeclare observability criteria to select a community-level site from a national weekly SARS-CoV-2 wastewater database, excluding model performance from selection. The task is strictly causal: a 12-week window predicts the next weekly log-transformed concentration; all preprocessing is fitted within training folds. Persistence, EWMA, Random Forest, XGBoost, LSTM, and GRU receive identical forecast-time information under a fixed hyperparameter budget, with matched missingness information provided across model families. Five disjoint chronological origins are evaluated on fixed held-out weeks with five seeds per stochastic model. Three degradation axes are then applied: density reduction at multiple calendar offsets, contiguous gaps of matched length at early, middle, and recent positions, and history truncation. These represent reduced sampling capacity, temporary monitoring interruption, and late entry into surveillance. Results. Under native conditions the local frontier is RMSE 15.70 (EWMA); the strongest learned model trails by 13\%. The frontier holder rotates across all five origins. Temporal variability dominates optimisation variability for trees (SD 5.0–5.9 vs < 0.3) but not for recurrent models, where one origin produces within-seed SD exceeding 20, revealing an optimisation-fragility mode absent from the tree family. Halving weekly density across the entire record raises the frontier by 1.08 (6.9\%), but gap placement produces a sharper effect: at 75\% retention, repositioning a 13-week gap from early or middle (15.69, 15.64) to recent (19.16) raises it by 22\% and shifts its holder from EWMA to XGBoost. Within this site, recent monitoring continuity carried forecast-relevant information that older observations could not replace — consistent with a signal whose effective properties shift with hydraulic and epidemic conditions. This has direct implications for programs resuming sampling after funding gaps or seasonal shutdowns. Under 52-week truncation, trees degrade substantially (XGBoost +11.18; RF +5.31) while recurrent response diverges (GRU −0.46; LSTM +2.94); these are mechanistically distinct failures: training-mass loss versus initialisation sensitivity, and neither family offers a robust short-history solution for communities recently joining monitoring networks. Implications and limitations. The reusable contribution is the evaluation protocol, not the site-specific numbers: the target-side reference varies with the observation regime in both level and model identity, so transfer evaluation requires scenario-specific comparison, not a single baseline. Any cross-site transfer must exceed local baselines predeclared before source data is introduced; source-side initialisation should additionally compress recurrent seed dispersion, yielding a falsifiable signature. Although the protocol is reusable, current limitations include single-site evaluation, synthetic rather than outcome-dependent missingness, retrospective model selection, and unresolved cross-site assay comparability.