Same Text, Different Prediction: Serving-Context Nondeterminism in Text Classifiers
Abstract
On-device and edge deployment runs text classifiers under reduced precision, integer quantization, and specialized runtimes chosen to meet compute, memory, and energy budgets, often with no cloud reference to check against. We show that these deployment settings, together with the batch a request happens to share, can change a classifier's prediction even when the trained parameters, the input text, and the tokenizer are fixed. Prior work documented such changes for text generation and attributed them to floating-point non-associativity and shape-dependent kernel selection; whether the same factors affect classification has had almost no direct measurement. To our knowledge this is the first systematic study of serving-context sensitivity in neural text classifiers: we train 180 models and evaluate them across 3,960 serving contexts spanning discriminative, pseudo-generative, and fully generative formulations. Label stability conceals score instability: changing only the batch shape flips no labels across 226,104 full-precision comparisons, yet under bf16 it redistributes up to 56.7 percentage points of predicted probability, with the flip rate reaching 30.6\% among examples whose top-two margin is below 0.01, and fully generative classifiers are the most exposed. We trace the numerical batch effect to shape-dependent matrix-multiplication reductions removed by batch-invariant operations, while assigning token positions by real-token order removes the larger padding effect. Reliable on-device classification therefore requires fixing these serving settings and monitoring class probabilities, not the predicted label alone.