When Depth Lies: Benchmarking Vision-Language Models on Mirror-Induced RGB-D Ambiguity
Abstract
Vision-language models (VLMs) have achieved significant progress in scene understanding and spatial reasoning. However, existing evaluation benchmarks predominantly assume that all visible content corresponds directly to physical entities in the scene, leaving model performance in reflective scenarios largely unexamined. Reflective surfaces preserve object appearance while altering the mapping between appearance and physical location, causing VLMs to produce systematic reasoning errors in such scenes. In this work, we introduce (i) a grounded RGB-D dataset for mirror-aware spatial reasoning, and (ii) the first comprehensive and systematic evaluation of state-of-the-art VLMs in reflective scenarios. Specifically, we collect 1,011 RGB-D images from 525 real-world scenes, spanning diverse indoor and outdoor environments and six reflective surface types, under varying lighting conditions across both daytime and nighttime settings. All images are annotated with reflective surface region masks and fine-grained, polygon-level instance annotations for both directly observed and mirror-reflected objects, with depth maps serving as both geometric ground truth for annotation and optional model input. We curate a mirror-induced reasoning ambiguity benchmark, named \textbf{MIRA-Bench} (Mirror-Induced Reasoning Ambiguity Benchmark). MIRA-Bench comprises 2,955 unique QA pairs across eight sub-tasks organized into three cognitive levels: \mbox{Reflection-aware} Perception, Spatial Reasoning, and \mbox{Scene-level} \mbox{Decision-making}, with five object-referencing protocols yielding 14,140 QA samples. We evaluate state-of-the-art VLMs on MIRA-Bench and find that these models consistently underperform across tasks, suggesting that current VLMs do not yet understand mirrors. Common failure patterns include misclassifying whether objects are directly observed or seen through reflections, failing to identify that reflected and directly observed instances correspond to the same physical object, producing erroneous cross-space spatial inferences, and tending to adopt conservative default strategies in action-oriented decisions instead of reasoning about the true scene layout. We will release the dataset and evaluation code to support future research.