Beyond Pattern Matching: Bridging the Sim-to-Real Gap in Transient Flare Verification Using Vision-Language Models
Abstract
Flares are sharp brightenings of astrophysical objects that constrain the properties of the underlying source, including quasars powered by accreting supermassive black holes. Flare searches end with human visual inspection of each candidate, the one stage that resists automation: real flares are too rare to label, and labels cannot be borrowed across surveys with different cadence, depth, and noise. We therefore simulate five classes (Gaussian, gamma, and fast-rise exponential-decay flares, plus quiescent baselines and single-epoch spikes) and test whether models trained on them find real flares in the light curves of spectroscopically confirmed SDSS Stripe 82 quasars. CNNs do not close the sim-to-real gap: ResNet-18s reach 70.6\% (from scratch) and 72.9\% (ImageNet-pretrained) five-class synthetic accuracy, but neither transfers. One collapses all 92 real candidates into a single class, and the other recovers 3 of 27 genuine flares, no better than chance. We benchmark 11 vision--language models zero-shot on the same synthetic test data; the best reaches only 42.8\%, yet three of them combined into an ensemble recover 16 of 27 real flares (59.3\% recall, 55.2\% precision), five times the best CNN, making them a practical verifier when no real labels exist.