TDBench: Benchmarking Vision Language Models on Top-Down Image Understanding
Abstract
Top-down images are critical in applications such as autonomous navigation and aerial surveillance, whereas Vision-Language Models (VLMs) are primarily trained and evaluated on front-view benchmarks, leaving their performance in top-down settings largely unexplored. Moreover, conventional evaluation protocols based on single-pass accuracy or repeated testing on identical questions can overestimate model capability due to hallucinations or chance correctness. To address these limitations, we introduce a 2{,}000-question benchmark over ten task categories for drone-altitude image understanding, and RotationalEval (RE), an evaluation protocol that leverages rotational invariance to measure answer consistency across multiple orientations of the same scene. Human evaluators lose almost nothing going from single-pass accuracy to RE, whereas across the 60 open-source and proprietary VLMs we evaluated, accuracy drops by 15 to 26 points. We analyze per-question rotation-outcome distributions across the models we evaluated with an Empirical Mixture Model that compares each model's distribution over the number of correctly-answered rotations (0 to 4) to a difficulty-adjusted Poisson-binomial null, splitting the residual into a mastery component and a consistent-failure component. The benchmark is integrated into an open-source evaluation toolkit and will be released upon publication.