Fidelity metrics for synthetic clinical time series invert under selection
Abstract
Synthetic clinical time series are accepted for use on the strength of fidelity metrics that compare them with real data. Those metrics are validated by showing that they separate real records from synthetic ones, which does not establish that they rank generators in the order a clinician would. We build a generator family carrying a structural failure whose cause is known and controllable, and score it under an acceptance battery before and after structures are selected on each metric. Averaged over the generator space, one widely used estimator carries no usable information about physiological plausibility, and it becomes strongly anti-correlated with plausibility as soon as structures are selected on it. Selecting the same number of structures at random has no such effect. The behaviour follows from how the estimator pairs rows, and it does not preserve a previously published comparison of generative models.