RaZeR: Pushing the Limits of NVFP4 Quantization with Redundant Zero Remapping
Abstract
The recently introduced 4-bit floating-point format, NVFP4, demonstrates remarkable performance and memory benefits for quantized large language model (LLM) inference. However, we observe two types of redundancy in the existing NVFP4 encoding: (1) The FP4 format naturally exposes an unused quantization value due to its sign-magnitude representation that contains both positive and negative zeros. (2) The FP8 block scaling factor, with 4-bit exponent and 3-bit mantissa, contains an unused sign bit given that the scaling factor is always positive. Additionally, we find that LLM weights are more tolerant to a lower-precision block scaling factor, such as 6 bits with 3-bit exponent and 3-bit mantissa. Based on these observations, we propose Redundant Zero Remapping (RaZeR), an enhanced numerical format that pushes the limits of NVFP4 for more accurate LLM quantization under the same memory footprint. RaZeR leverages the unused bits in the block scaling factor to adaptively remap the negative FP4 zero to a set of pre-defined special values, which maximally utilizes the NVFP4 encoding and better fits the LLM tensor distribution. Extensive experiments validate RaZeR’s superior performance for 4-bit LLM quantization. For example, RaZeR reduces the average perplexity loss of NVFP4 by 29.4% and 32.3% under weight-only and weight-activation quantization, respectively.