Where to Predict in Retinal OCT? Anatomy-Guided Target Selection for I-JEPA
Abstract
Learning transferable representations from retinal optical coherence tomography (OCT) is challenging because disease-related changes can be subtle, and large, consistently labelled datasets require substantial expert annotation. Self-supervised learning offers a way to use unlabelled scans and reduce dependence on diagnostic labels. Image-based Joint-Embedding Predictive Architecture (I-JEPA) learns by predicting target-region representations from visible context rather than reconstructing pixels. However, its uniform target placement does not account for the layered organization of retinal OCT, where retinal tissue occupies only part of each B-scan. We investigate anatomy-guided target selection using image intensity and predicted retinal segmentation while retaining the I-JEPA architecture and latent prediction objective. We compare tissue-directed rectangular targets, anatomy-shaped targets, and coverage-constrained placement with uniform masking, and evaluate the learned representations through frozen-encoder glaucoma classification on FairVision. Tissue-directed rectangle placement improves the observed classification results, including with segmentation-free intensity guidance, while more anatomically specific strategies do not consistently provide further gains. Analysis of the actual masks shows that target selection, target size, and visible context are coupled. Our findings motivate anatomy-guided predictive learning for retinal OCT and better-controlled comparisons of what to predict and what to keep visible.