A Coherent Encoder for Q2\_0, and the Limits of Free Permutations
Grant C Forbes ⋅ Sriman V Donthireddi ⋅ Jonathon Romero ⋅ Zixiang Tang ⋅ Adrian Chan ⋅ Jianbang Zhang
Abstract
The 2-bit methods for LLM weight quantization that work best all change the inference kernel --- codebooks, trellises, online rotations --- raising inference cost and divorcing training-time fake-quantization from deployed arithmetic. The legacy \texttt{Q2\_0} format needs none of that and already ships, but no published encoder makes it usable. We give one. A per-block scale search and a signed scale, combined with standard GPTQ and channel scaling, turn a model degenerate under the reference encoder into one that generates coherent text, and that is the difference between an initialization distillation can repair and one it cannot: after an identical recipe the reference encoder's model reaches a benchmark mean of only $6.4$ and still scores zero on five of eight benchmarks, while ours reaches $26.9$. We then ask whether \emph{block membership} can be chosen better by permuting the axes a transformer leaves free. We prove that the transformations of a gated block's intermediate axis absorbable into its weights are exactly the monomial matrices, so permutation is the whole of the discrete freedom --- and find it buys about one point of benchmark mean. We trace this to structure rather than search: one permutation must serve every row of a tensor, so both exactly-free relayouts in a transformer fail for the same reason, locating the remaining headroom in relaxations that break the sharing at a cost.
Chat is not available.
Successful Page Load