Simulator Brightness Shortcuts and Their Effects on AUROC: A Case Study in Strong-Lens Finding
Wasiq Amir
Abstract
When a simulator supplies training positives against real negatives, a classifier can exploit a simulator-side artifact that is not the property of interest. Global ranking metrics may still look strong while sensitivity at the operating threshold that matters for deployment stays weak. We diagnose this failure mode in a recent sim-to-real strong-lens finding study (Parul et al., 2025). The case is a brightness confound between simulated lenses and real non-lenses (Cohen's $d = -4.2$), consistent with a magnitude-boosting step applied only to the simulated class. A linear model using only 8 brightness statistics, with no spatial information, recovers 95--102% of the paper's naive AUROC under true generalization while capturing under 20% of TPR at 1% FPR on that same protocol. A ResNet under a matched protocol shows bulk confidence--brightness coupling ($\rho = -0.80$) that flips sign in the pooled high-confidence tier ($\rho = +0.16$); that positive tier correlation is concentrated in non-lenses ($\rho=+0.199$), while tier lenses show no significant relationship ($N=33$; CI includes zero). Rank-offset windows place the pooled flip in the high-score region, not at a special 1%-FPR cut. Across three training seeds with fixed data split, brightness decorrelation raises TPR@1%FPR in every run (mean $\Delta$TPR $= -0.063\pm0.033$), while AUROC $\Delta$ is unstable ($+0.014\pm0.053$). The empirical gap remains: AUROC alone can look strong while TPR at low FPR stays weak. We recommend reporting TPR at low FPR alongside AUROC, and testing brightness decorrelation across seeds, in future sim-to-real lens-finding work.
Chat is not available.
Successful Page Load