Stranger Things: When Objects Appear Without Their Typical Neighbours
Abstract
Visual context is a powerful cue for object recognition: scenes inform what, where, and at what scale objects are likely to appear. Yet over-reliance on context becomes a failure mode when models predict objects based on their typical neighbours rather than the objects' own visual features, a phenomenon we refer to as typicality bias. While this issue has been extensively studied in the pre-Transformer era, it has often been assumed that architectural and scale improvements would naturally solve it. Despite recent attempts in isolated settings, a systematic diagnostic across the modern vision stack remains absent. In this work, we provide such a diagnostic framework. Specifically, we measure typicality bias via normalized pointwise mutual information (NPMI) over object pairs, and introduce STRANGE-Bench (Synthetic Typicality-RANked Grounded Evaluation Benchmark): a photorealistic, layout-controlled benchmark spanning 11 typicality tiers from typical to never-co-occurring object pairs. Across \five model families (specialist detectors, open-vocabulary detectors, self-supervised, vision-language, and multimodal large language models) and three recognition tasks}(supervised detection, open-vocabulary detection, and classification), we find the bias to be highly prevalent: gaps of 5--11\% persist when objects appear outside their typical company, and they do not consistently shrink with scale. This persistence motivates a complementary question: how to mitigate it? Data augmentation is a natural lever, since it directly shapes the training distribution that produces the bias; yet existing augmentations are designed for visual diversity, not for breaking co-occurrence shortcuts as they sample from the very distribution that produces typicality bias, preserving rather than perturbing it. We therefore propose Anchor-Atypicality Sampling (AAS), a simple context-aware variant of paste-based augmentations that uses the NPMI signal to deliberately tilt the training-time co-occurrence distribution toward atypical partners. AAS reduces the typicality gap by up to 3.3 pp on natural-image COCO splits and 17.4 pp on synthetic STRANGE-Bench across three representative model families (specialist detection, open-vocabulary detection, self-supervised classification), without sacrificing overall performance.