Error Diffusion Replaces Latent Weights: Training Binary-Weight Language Models on a Few Bits of Dithered State
Abstract
Binary-weight language models store one bit per weight, but training them keeps a full-precision latent copy of each weight plus Adam state, about 96 extra bits per weight, so a deployed binary model cannot be adapted on the device that runs it. We replace the latent weight with a 2-8-bit residual buffer quantized by stochastic rounding, which cuts the training state per weight from 97 bits to 3-9, up to 19x less. The resulting error-diffusion (ED) optimizer moves the residual-carrying loop of image halftoning onto the weight-update path. Each buffer accumulates the unrealized part of every update and flips the weight when the accumulated value crosses a threshold, and it works only when its quantization grid has a level exactly at that threshold. Under this threshold-alignment rule, 4-bit buffers match 8-bit buffers on a binary multilayer perceptron benchmark, and the same rule lifts a different latent-free optimizer, Bop, from a 10-point deficit at 4 bits to full-precision accuracy. On character-level transformers with 97% of parameters binary, ED trains within 5-11% of the latent-weight reference at a 2,000-step budget and within 1.7% at convergence. On scaling ladders over enwik8 and wikitext-103 up to 19M parameters, the latent-free loss stays within -2.5% to +4% of the reference at five of six sizes and trails by 10% at the largest tokenized size. A deployed binary model continues training under ED at one tenth of the latent training state, matching latent-weight continuation, and a fused CPU kernel completes the adaptation within a 275 MB training-memory budget, which latent-weight training cannot. A 250-step probe sets the flip learning rate at any new scale.