Eliminating Nondeterminism in LLM Inference via Integer-Only Arithmetic
Abstract
LLM inference is nondeterministic even at temperature 0. Floating-point addition is not associative, so the logits depend on the reduction order of whichever kernel runs, and batch shape, prefill vs. decode, library heuristics, and hardware all change which kernel that is. Batch-invariant kernels address this by pinning a single order, which scopes determinism to those kernels on that platform. Integer addition, by contrast, is associative: every reduction order yields the same bits. DetLLM runs the entire forward pass of Qwen3-0.6B (every matmul, softmax, RMSNorm, SwiGLU, and rotary embedding) in exact integer arithmetic, with no floating-point operation anywhere between input ids and int32 logits. Nine 512-token greedy generations spanning NVIDIA A100 and H100 GPUs and AMD, Intel, and Apple CPUs (four vendors, three instruction sets) produce one identical chained hash over all 512 logit vectors. Nine matched fp16 runs produce nine distinct hashes, and every pair disagrees from the first token. One int8 artifact serves every device, server GPU or laptop CPU, with nothing tuned per platform. WikiText2 perplexity comes out to 20.72 vs. 20.95 for fp16. Compiled, CUDA-graphed integer decode reaches 106 tok/s at batch 1 on an A100, 3.6× an eager fp16 baseline; we did not measure a comparably optimized fp16 baseline, and prefill runs at 0.13× fp16.