V-LUMEN: Visual Lookup Memory for Embedding Scaling in Vision-Language Models
Abstract
Embedding scaling has recently emerged as a promising direction for increasing the capacity of language models through embedding-level representations, rather than solely expanding transformer depth or width. However, existing embedding scaling methods are primarily designed for discrete text tokens, leaving it unclear whether similar mechanisms can be extended to vision-language models, where visual representations are continuous, high-dimensional, and spatially structured. In this paper, we introduce V-LUMEN, a visual lookup memory framework for embedding-level scaling in vision-language models. V-LUMEN augments visual-language representations with reusable visual embeddings retrieved from an external memory indexed by discretized visual patterns, thereby expanding accessible representational capacity without directly increasing the dense computation of the backbone model. To address the unique challenges of visual representations, V-LUMEN introduces spatial aggregation for constructing structured visual memory keys and text-conditioned hashing for retrieving task-relevant visual memories based on the input query. Together, these components enable memory retrieval that is both spatially aware and instruction-conditioned. Experiments across diverse reasoning-intensive multimodal benchmarks show that V-LUMEN effectively augments vision-language models through external visual memory. Without backbone adaptation, V-LUMEN substantially improves the vanilla Qwen3-VL-2B-Instruct model, and further achieves competitive or better performance against parameter-matched MoE and LoRA baselines, demonstrating the potential of visual embedding scaling as an alternative scaling axis for vision-language models.