Verifier Choice is a Benchmark Design Variable: Auditing Structural Counting Evaluation in Text-to-Image Models
Shurun Li ⋅ Michael H Wang ⋅ Haibo Zeng ⋅ Long Wang
Abstract
Automatic evaluation is now central to text-to-image benchmarking, but the verifier that converts generated images into scores is often treated as an implementation detail. We argue that this is unsafe: an automatic benchmark score is a property of the image generator, prompt distribution, verifier, verifier-use protocol, and scoring rule, not of the generator alone. We study this issue in structural counting, a controlled setting where human-visible counts can be audited and where both instance-level object counts and parent-bound substructure counts arise naturally. We construct a verifier-centric audit suite over seven counting objects: cell phones, chairs, and traffic lights for object counting, and clover leaflets, dresser drawer fronts, power-strip sockets, and turbine blades for substructure counting. Across 6,240 generated images from Qwen-Image-2512 and FLUX.2-dev, we compare open MLLM verifiers under count-only, reject-aware, zero-shot and few-shot protocols, plus TIFA-like target-query variants and targeted frontier closed-model sanity checks on two hard objects. We also compare MLLMs with a detector baseline. On a balanced 2,280-image evaluation split, Qwen3.5-27B with few-shot reject/count achieves the best image-level agreement among open verifiers ($\kappa=0.754$) with a small aggregate score gap ($+1.54\text{pp}$), but changing the protocol, model size, or generator distribution can shift automatic scores by tens of points. Targeted GPT-5.5 and Claude Opus 4.7 checks do not remove the failure modes: GPT-5.5 still has only 59.92\% clover valid-count fidelity, while Claude Opus 4.7 has lower same-slice agreement ($\kappa=0.432$). We propose reporting a factorized verifier reliability profile---aggregate score gap, image-level agreement, valid-image count fidelity, rejection behavior, output compliance, and directional bias---rather than a single automatic benchmark score.
Chat is not available.
Successful Page Load