Whole-Model Quantization Below 4 Bits Per Weight for On-Device Multimodal Models
Abstract
On-device deployment is bounded by model footprint, and moving capable mul-1 timodal models onto phones means compressing the whole model: existing post-2 training methods degrade below 4 bits and leave the per-layer embedding tables,3 modality towers, and vision/audio bridges of a multimodal model unaddressed en-4 tirely. We present CQ1, a single rotation-and-codebook recipe applied uniformly5 to every weight tensor. CQ folds a SmoothQuant-style per-channel input scale into6 the weights, applies a deterministic structured Hadamard rotation shared across7 all tensors and models, L2-normalizes each group, and assigns a single shared8 Lloyd–Max codebook trained on the unit-norm Gaussian, with an optional block-9 wise GPTQ residual pass for tensors that have an activation Hessian. It extends the10 rotation-based formalism of TurboQuant from KV-cache activations to all offline11 weight tensors and drops the runtime QJL bias-correction stage. In Gemma-4-12 E2B, the per-layer-input tables alone hold 52% of parameters and are untouched13 by linear-only quantizers; whole-model coverage is what makes deployment bun-14 dles below 4 bits per weight reachable at all. Across Gemma-4-E2B, Qwen3-15 VL-2B, and LFM-2.5-VL, CQ is near-lossless on text and function calling at 3.2616 average bits, leads the scalar PTQ baselines we evaluate (GPTQ, AWQ, HQQ)17 at every setting below 4 bits, and its production bundle shrinks Gemma-4-E2B18 from 4.79 GB to 0.9 GB while matching or exceeding matched-bpw GGUF de-19 ployments on most metrics.