Q-JEPA: Joint-Embedding Predictive Adaptation for Quantized Large Language Models
Abstract
Low-rank adaptation and low-bit quantization make large language model fine-tuning substantially more memory efficient. Recent work on LLM-JEPA shows that next-token prediction can be improved by incorporating a joint-embedding predictive objective over paired representations. The interaction between this representation-space supervision and low-bit adaptation remains largely unexplored. We introduce Q-JEPA, a controlled objective--precision study that compares next-token and JEPA-augmented training under BF16 LoRA and 4-bit NF4 QLoRA with matched optimization. We evaluate Q-JEPA using Llama-3.2-1B-Instruct on NL-RX-SYNTH, a natural-language-to-regular-expression benchmark. Across all matched fine-tuning seeds, incorporating JEPA improves exact-match performance under both LoRA and QLoRA. The improvement remains consistently positive under NF4 adaptation, although its absolute magnitude is smaller than under BF16 LoRA, showing that adaptation precision modulates the benefit of joint-embedding supervision. Overall, our results establish that predictive representation learning remains complementary to quantized low-rank adaptation rather than being specific to higher-precision fine-tuning. Q-JEPA provides a controlled framework for studying how representation-space objectives interact with numerical precision and motivates memory-efficient adaptation methods that better preserve their benefits under quantization.