Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context
Abstract
Long context is becoming a core capability of modern large vision-language models (LVLMs), enabling sustained context management across long-document understanding, video analysis, and multi-turn tool use in agentic workflows. However, practical training recipes remain under-documented, such as how to construct and mix long-context data. In this work, we present a systematic study of long-context continued pre-training for LVLMs, extending a 7B LVLM from 32K to 128K context with extensive ablations on long-document data. We find that: 1) synthesizing long-document VQA data provides effective and diverse long-context supervision, covering tasks from information extraction to numerical reasoning; 2) retrieving relevant evidence remains the primary long-context bottleneck, favoring retrieval-heavy mixtures with a small amount of reasoning data to preserve task diversity; 3) surprisingly, pure long-document VQA data largely preserves short-context capabilities, suggesting that instruction-formatted long data lessens the need for short-data mixing. Instantiating these findings, we obtain MMProLong by continuing training from Qwen2.5-VL-7B, improving long-document VQA performance by 7.11 points under only a 5B-token budget. More importantly, MMProLong generalizes beyond its 128K training window, maintaining strong performance at 256K and 512K without additional training. More broadly, it also transfers to webpage-based multimodal needle-in-a-haystack tasks, long-context vision-text compression, and long-video understanding without task-specific supervision. Overall, our study provides a practical LongPT recipe and an empirical foundation for advancing the next generation of LVLMs.