BlackLight: Differential Testing on Inference-Time Deep Learning Systems
Abstract
Robustness measured during training is assumed to carry over to deployment. Yet what ships is a compiled, inference-only binary, not the validated checkpoint. Quantization, operator fusion, and hardware lowering make its predictions diverge. Host-side validation measures the checkpoint, not the artifact that ships. Sampling clean inputs on the device rarely surfaces it, and on some pairs never. Deployment platforms lack automatic differentiation. That renders gradient-based testing infeasible and forces existing approaches onto inefficient brute-force search. We present BlackLight, a controlled differential testing framework for auditing a deployment pair before release. It makes the most of the two signals an owner already holds: the original model’s gradients and the probabilities the device returns. Together they narrow the ambiguity a blind search must otherwise resolve by sampling. A two-term objective holds the original model correct while maximizing prediction divergence. We evaluate on real MCU, TPU, DSP, and GPU hardware across image, audio, sensor, and malware tasks. BlackLight surfaces a disagreement on over 90% of random seeds in nearly all supported settings. Held to the same perturbation budget, black-box and fuzzing baselines reach 30% at best and 0–6% on average, and they are given 180 s per seed where BlackLight needs 2.4–4.1 s. A deployment pair ships to an entire fleet, and so do its hidden failures. An owner should audit the artifact that ships, before it ships.