SANEval: Open-Vocabulary Compositional Benchmarks with Failure-mode Diagnosis
Rishav Pramanik ⋅ Ian Nielsen ⋅ Jeffrey Smith ⋅ Saurav Pandit ⋅ Ravi P Ramachandran ⋅ Zhaozheng Yin
Abstract
Text-to-image (T2I) models still fail on compositional prompts that require multiple objects, correctly bound attributes, accurate counts, or specific spatial relations, yet the benchmarks used to measure this progress are themselves limited: they rely on fixed-vocabulary detectors (typically the **80** MS-COCO classes), produce a single opaque score, and are rarely validated against human judgment. We introduce **SANEval** (Spatial, Attribute, and Numeracy Evaluation), an open-vocabulary compositional benchmark with two contributions: (i) a public dataset of $\sim$**5,250** prompts paired with images from six state-of-the-art T2I models, partitioned into a Simple split (algorithmically generated, human-validated) and a Hard split (fully human-written), and labeled by compositional task type (spatial, numeracy, color, shape, texture), and (ii) a modular evaluation pipeline that combines an LLM-based prompt parser, an open-vocabulary detector with LLM-driven synonym expansion and mapping, and three scorers for attribute binding, spatial relations, and numeracy. Each scorer emits both a numerical score and structured diagnostic feedback identifying what is missing, extra, or mis-bound. We validate SANEval against human ratings ($n=500$, **4** annotators, **5** categories): SANEval achieves a positive Pearson correlation with humans in every category (avg $r=0.222$), with **4** of **5** category-level correlations significantly above zero (Fisher **95%** CIs excluding **0**); the strongest prior compositional benchmark averages near zero ($r=-0.016$), with **4** of **5** category CIs spanning zero. SANEval is, to our knowledge, the only compositional T2I benchmark whose correlation sign agrees with humans on every axis. A manual audit of **10,040** detections decomposes the pipeline into a perception stage and an LLM/VLM verification stage and measures end-to-end precision at **96.6%**; the verification stage (synonym mapping plus VLM-as-judge attribute reading) contributes only **1.05%** of audited false positives, while removing it collapses scores by **43–76%** in ablation — substantial detection work at modest precision cost. The dataset is publicly released at [https://huggingface.co/datasets/saneval-ann/saneval-release](https://huggingface.co/datasets/saneval-ann/saneval-release) and the evaluation pipeline at [https://anonymous.4open.science/r/saneval-anon](https://anonymous.4open.science/r/saneval-anon).
Chat is not available.
Successful Page Load