Same Command, Different Answer: vLLM’s Runtime LoRA Path Breaks Within-Configuration Reproducibility
Abstract
That greedy LLM decoding is not bit-reproducible is well documented, and the accepted account is numerical: floating-point non-associativity combined with kernels whose reduction order varies with batch size, GPU count, tensor-parallel degree, or precision. The folk workaround follows: serve at batch size one. We quantify a source that survives it. Inside a single vLLM process at concurrency one, with caching off and everything else fixed, we vary only which model a request names: requests to the base model are byte-identical, requests naming a LoRA adapter are not. Across 8 independent server processes the base emitted one distinct output hash in 24 runs and the adapter emitted twenty-four; adapter agreement averages 0.191 (bootstrap over sessions, [0.156, 0.222]). Running three replicates rather than two separates three mechanisms that a single pair conflates: concurrency degrades the base model as expected, caching produces a transient confined to a session’s first run, and the adapter path shows neither: it fails with one adapter alone in the server. Excluding that transient, all four configurations that do not activate an adapter reproduce exactly at concurrency one; both that do, do not. Of three determinism controls the server already ships, two change nothing, including the one a 2024 bug report names as the fix, while batch-invariant kernels restore byte-identical output completely, at about 3× the cost. That this works at concurrency one, where there is no batch variation left to make invariant, shows request-batch composition is not a necessary trigger and narrows the fault to the adapter-serving path itself, which we do not claim to localise further. In an eight-agent pipeline the disagreement compounds until a fifth of binary pass/fail verdicts flip between identical runs (a tenth once answers are allowed to finish rather than truncate), and one method’s spread across reruns exceeds the gap between the two methods compared by eight times. Replaying it through in-process adapter switching gives zero divergences over 1280 paired comparisons, at a fifth of the serial throughput.