Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs
Luka Ribar ⋅ Jeevan Bhoot ⋅ Douglas Orr
Abstract
Deploying vision-language models (VLMs) on mobile devices is challenging due to their significant memory and compute requirements. Our method for quantizing VLMs for efficient inference on resource-constrained hardware combines a quantization-aware training (QAT) pipeline that uses the model itself to generate training data with a novel $2.7$-bit-per-parameter format and $8$-bit activations, supporting efficient execution on Arm CPUs. We validate our approach by compressing the Llama 3.2 11B Vision Instruct model to $3.7$ GB, preserving strong performance on a set of standard visual question answering tasks.
Chat is not available.
Successful Page Load