Is Self-Pretraining Really Useful for Diagnosis in Medical Time Series?
Abstract
Headline accuracy numbers are the currency in which machine learning capability is communicated to clinicians, reviewers and regulators, yet they are quietly conditional on training choices that are rarely reported. We study one such choice: whether a transformer is trained from scratch or initialised by Self-PreTraining (SPT), i.e.\ masked reconstruction on the target dataset itself, with no external corpus. Across three clinical time-series tasks locomotion classification (CAMARGO), wearable stress detection (Non-EEG) and Parkinson's disease detection from gait (Gait PD) and four masking objectives (point, block, column, mixed), we compare identical architectures at depths of one to three layers. SPT never degrades performance and improves accuracy by up to 9.9 percentage points, with a mean multivariate gain of 2.1-4.7 points depending on masking strategy and depth; gains widen as models get deeper. Two consequences follow. First, SPT is a cheap, data-local and architecture-agnostic way to make transformers usable in the small-cohort regime typical of clinical work. Second, and less comfortably, an accuracy difference of this magnitude is larger than the margins by which competing medical models are routinely declared superior which means a substantial part of what is communicated as model capability is attributable to initialisation, an implementation detail that most papers, press releases and clinician-facing summaries omit entirely. We argue that initialisation belongs in the reporting minimum for biomedical ML, and give a concrete disclosure template.