Not Every Image Teaches Vision: Visual-Necessity-Gated Continual Learning for Multimodal Large Language Models
Abstract
Continual tuning is essential for adapting multimodal large language models to evolving tasks and domains, yet existing methods often assume that every image-present example provides valid supervision for the visual pathway. We address the resulting gap: how to learn from all incoming multimodal instructions while preventing visually unnecessary supervision from drifting the visual interface. We propose Visual-Necessity-Gated Continual Tuning (VNG-CT), a path-aware framework that estimates sample-level dependence on visual evidence by comparing target likelihoods under the original image and a counterfactual null image. The resulting gate routes gradients to different parameter paths: all examples update the language path, visually necessary examples update the visual path, and low-necessity examples are absorbed by a lightweight calibration path that promotes invariance to irrelevant visual evidence. Across CoIN, MLLM-CL, and UCIT-style continual instruction streams, VNG-CT improves final and average performance while reducing visual forgetting. These results suggest that visual necessity offers a practical principle for stabilizing continual multimodal adaptation without discarding useful language-side supervision.