Virtual Task Prompting for Multi-Task Scene Understanding
Abstract
Multi-task scene understanding models typically require modifications to the vision encoder or complex decoders for cross-task interaction, which limits their flexibility and compatibility with off-the-shelf vision foundation models (VFMs). This paper introduces Virtual Task Prompting (VTP), which frames multi-task prompting as retrieving task-adaptive subsets from a shared virtual state, where both the state and accessing process itself are optimizable in an end-to-end manner. Concretely, VTP maintains a set of virtual tasks alongside the model, and at each layer, learnable routing matrices read an actual task composition as input for interaction with image tokens, then write the processed results back, forming explicit learning pathways that capture the intricacies among tasks before and after the computation of each encoder block. Since VTP keeps the backbone clean, it integrates seamlessly with recent VFMs like DINOv3, advancing the state of the art on Pascal-Context and NYUD-v2. Stress test on Taskonomy with 13 heterogeneous tasks further shows VTP's robustness in extreme settings. Ablations and analysis confirm that VTP's gains stem mainly from its prompting design rather than merely increased capacity.