From Preprocessing to Downstream Validity: Auditing Transformed Targets in Multimodal Health AI
Abstract
Shared representation spaces and transformed targets are becoming important parts of scientific and clinical health-AI inference systems. We propose that, by itself, successful preprocessing does not provide enough evidence that a transformed target is valid for the downstream claim made using it. For instance, we analyze a cross-modal fMRI decoding pipeline using ImageBind representations, where audio embeddings were residualized with respect to video before evaluation. True video could still retrieve the residualized audio targets well above chance under the same downstream retrieval metric. This remained the case with validation-selected ridge and with a nonlinear residualizer that achieved stronger held-out nuisance prediction. However, this does not demonstrate a failure in a clinical system, nor does it imply that residualization is generally invalid. Instead, it highlights an additional difficulty in evaluating health AI: a transformation may achieve its goal based on the criteria used to define it but fail to produce a valid target for the subsequent interpretation. Consequently, we suggest testing whether an identifiable nuisance source can still access the transformed target under the same representation space and evaluation metric used to support the scientific or clinical claim. If the alternative path is still open, then the downstream result by itself is not enough to support the interpretation we intended to test.