NyoomFloat12: Accelerating LLM Inference via Lossless 12-bit Weight Compression
Sylvie Liberman ⋅ Xinyu Fang ⋅ Tianyi Zhang ⋅ Tri Dao ⋅ Dan Fu
Abstract
BFloat16 weights of large language models carry only ${\sim}10.5$ bits of information per 16-bit value, yet exploiting this redundancy on GPUs, where weights bottleneck both per-token bandwidth and concurrent capacity, has not produced wall-clock speedup. Existing variable-length lossless codes are compute-bound on SIMT decode. We present NyoomFloat12 (NF12), a lossless 12-bit fixed-length format for BF16 weights that addresses both bottlenecks. Over $99.7\%$ of trained BF16 weights have the upper four exponent bits set to $\texttt{0111}$ after a per-matrix power-of-two scale; dropping them yields a fixed-length 12-bit format with 16-instruction SIMT-friendly decode. Out-of-range groups reuse the encoded slot to index a per-matrix verbatim BF16 buffer at zero per-group storage overhead. In the bandwidth-bound regime, fused NF12 GEMV achieves up to $1.21\times$ matched-BF16 and $1.35\times$ cuBLAS, $71\%$ of the entropy-bound optimum (a hypothetical implementation reading the joint entropy floor at full HBM bandwidth). In the capacity-bound regime on Qwen3-32B, NF12's smaller footprint opens $4.5\times$ more KV-cache room; with decode at near-peak HBM bandwidth ($3.3\times$ faster than prior lossless decoders), NF12 raises throughput up to $2.67\times$.
Chat is not available.
Successful Page Load