Graph-Supervised Adaptation of CLIP Improves Relational Compositionality by Reshaping Language Embedding Geometry
Nimay Upen Shah ⋅ Amit Sethi
Abstract
Contrastive vision-language models such as CLIP suffer from a precise and measurable failure in relational compositionality, a linguistic property governed by thematic role assignment. Sentence pairs sharing the same predicate but differing only in argument role ("a dog chases a cat" vs. "a cat chases a dog") collapse to nearly identical embeddings: in CLIP ViT-B/32 the mean cosine similarity among such pairs is $0.971$, compared with $0.7663$ across unrelated sentences. We attribute this failure to the EOS-pooling bottleneck in CLIP's text encoder combined with a training objective that optimizes only global image-text alignment, leaving no direct gradient pressure to preserve predicate-argument structure. We introduce GS-CLIP, a graph-supervised adaptation that injects relational computation into the text encoder itself, rather than treating structure only as a property of the training data: scene graphs extracted from captions are processed by a three-layer Graph Attention Network whose representations are fused into token-level hidden states via cross-attention, with LoRA adapters providing parameter-efficient adaptation. Training proceeds in two stages, Visual Genome scene graphs first associate relational language with visual evidence, then MS-COCO fine-tuning with a composite objective directly targets relational embedding geometry while preserving global alignment. GS-CLIP reduces intra-predicate cosine similarity from $0.971$ to $0.379$, $$ \Delta_{\text{IPS}} = \frac{0.971 - 0.379}{0.971} \approx 61\%, $$ a reduction in relational embedding collapse that comes with improved embedding isotropy, and translates to a $29.05$ percentage-point gain on the ARO benchmark and consistent gains on text-only relational discrimination tasks. Image-grounded compositionality beyond the evaluated benchmarks remains an open direction, and a preliminary, small-scale component ablation indicates the contribution of each architectural piece is not yet fully settled.
Chat is not available.
Successful Page Load