Patch4Patch: Restoring Structural Connectivity in Patch-based Vision Encoders
Abstract
As Multimodal Large Language Models (MLLMs) continue to advance and find increasingly broad applications, a key challenge persists: the patch-based vision paradigms widely adopted by MLLMs inevitably introduce structural fragmentation when processing information-dense images such as tables and charts. Specifically, when a table row or chart axis spans multiple visual partitions, the lack of cross-partition communication causes the model to lose track of structural continuity, leading to performance fluctuations when layout changes --- a problem we term layout sensitivity. To address this fundamental limitation, we propose Patch-for-Patch (P4P), a lightweight module that can be plugged into existing patch-based vision encoders. P4P dynamically selects a sparse set of tokens as bridging anchors and uses them to establish cross-partition communication through bidirectional attention. Combined with a gated fusion strategy, P4P restores global structural coherence without the quadratic cost of global attention. Extensive experiments across Qwen2.5-VL, MiniCPM-V, and InternVL3.5 demonstrate that P4P consistently improves structured content visual understanding, notably boosting TableEval accuracy by up to 9.12\% and significantly reducing layout sensitivity. The source code is available at: https://anonymous.4open.science/r/P4P/.