When Image-Based Metadata Imputation Adds Little: Evidence from Multimodal Skin Lesion Classification
Abstract
Clinical metadata may improve image-based diagnosis but is often incomplete where health records or digital infrastructure are limited. We test whether age group, sex, and anatomical site can be recovered from dermoscopic images and improve skin-lesion classification. Linear probes use intermediate representations from the frozen DermLIP-PanDerm model; cross-attention then combines the completed metadata with a ConvNeXt-Swin encoder. On a fixed image-level split of HAM10000, the probes outperform majority prediction and a diagnosis-trained ConvNeXt: AUROC 0.799 for sex, quadratic weighted kappa 0.577 for age group, and macro-F1 0.494 for anatomical site. The best multimodal classifier reaches macro-F1 0.8178, compared with 0.8074 for the image-only model. Up to 75% synthetic missingness, the three handling strategies differ by less than one macro-F1 point. Cross-modal imputation performs better in an image-free downstream ablation. These single-run results suggest that recovered attributes add little when the source image remains available, but they do not establish equivalence between strategies.