Multimodality as Supervision: Self-Supervised Specialization to the Test Environment via Multimodality
Kunal Pratap Singh ⋅ Ali Garjani ⋅ Rishubh Singh ⋅ Muhammad Uzair Khattak ⋅ Efe Tarhan ⋅ Jason Toskov ⋅ Andrei Atanov ⋅ Oğuzhan F Kar ⋅ Amir Zamir
Abstract
$\textbf{Cross-modal learning}$, i.e., learning to predict one modality from another, leverages multimodality as a source of self-supervision. Many practical applications, such as deploying a household robot, involve devices equipped with a rich set of sensors that can collect multimodal data in their deployment environment. This presents an opportunity to learn representations from this multimodal data through cross-modal learning, which is currently underutilized. Prior work has studied cross-modal learning on pre-collected internet-based datasets. However, these datasets provide only a fixed set of observations and no access to the underlying physical environments from which additional viewpoints or modalities could be acquired. We study a setup where a multisensory agent moves through the physical world and collects multimodal data. This gives us control over multimodal data collection, including the agent's operating space, viewpoints, amount of data, and available modalities. Within this setup, a particularly interesting instantiation arises when we restrict the agent's world to a single physical space. We call this the test space, and perform both self-supervised pre-training and evaluation within it. This results in a $\textbf{specialization}$ setup in which the representation is developed for a specific test space. It also reflects practical applications in which an agent repeatedly operates in a fixed physical space (e.g., household robotics). We call this Test-Space Training ($\texttt{TST}$). We find that cross-modal learning on rich multimodal test-space data enables \tsp to achieve state-of-the-art performance in that space, outperforming generalist baselines pre-trained on large-scale internet-based datasets. This reduces the need for external internet-scale pre-training datasets where the deployment environment is known.
Chat is not available.
Successful Page Load