Injection-based evaluation of binary-asteroid wobble in Gaia astrometry
Abstract
A binary asteroid is an asteroid with a satellite. As the satellite orbits, the system's position as measured by Gaia can wobble by up to a few milliarcseconds. Established discoveries tell us which objects are binaries, not whether the wobble leaves an observable pattern in the data, and the wobble is usually smaller than the error on a single measurement. We therefore simulate such signals ourselves, for training and for evaluation. Starting with a real object, either a known non-binary or an unlabeled object, which we call the carrier, we inject a synthetic wobble into its data. We then check whether a detector responds to the injected signal. Judging detectors on known binaries instead is biased. We find that models which separate known binaries from other objects do so largely by brightness, not by the wobble. A classifier that never sees the time series reaches ROC AUC 0.938 on that task. We propose a paired injection test that compares each carrier with its injected copy, so any score change is attributable to the injection. In the second-highest injected-SNR decile, a convolutional network on the time series ranks the injected copy above its twin in 80% of pairs (chance 50%), while a baseline on summary statistics stays below 54%. For scale, even a matched filter given the true waveform detects only 1.6% of these injections at a 1% false-alarm rate, so the paired test shows sensitivity where detection is out of reach. We then score ten binaries announced after our label freeze against 14,999 unlabeled objects. Models trained without unlabeled carriers rank them high, whereas models trained with them rank them low, regardless of architecture. We interpret this as domain shift, since the carriers determine which populations the model sees during training; neither behaviour certifies detection. Injection evaluation rests on two open conditions, that the simulated signals resemble real binaries and that the carriers resemble the objects searched. The benchmark it replaces can test neither.