VTBench: Disentangled and Human-Aligned Evaluation for Image-Based Virtual Try-on
Abstract
While virtual try-on has achieved significant progress, evaluating these models towards real-world scenarios remains a challenge: unpaired metrics (FID/KID) miss task-specific dimensions such as texture and consistency, and most test sets are limited to indoor scenarios. To address these needs, we introduce the Virtual Try-on Benchmark (VTBench), the first-ever hierarchical try-on benchmark suite that systematically decomposes virtual image try-on into hierarchical, disentangled dimensions, each equipped with tailored test sets and evaluation criteria. VTBench exhibits three key advantages: 1) Multi-Dimensional Evaluation Framework: The benchmark defines three top-level dimensions (General Image Quality, Garment Preservation, Auxiliary Consistency) and six fine-grained dimensions (Distributional Fidelity, Aesthetics, Texture Fidelity, Cross-Category Plausibility, Background Consistency, and Hand Consistency). Granular evaluation metrics and targeted test sets pinpoint model capabilities and limitations across diverse, challenging scenarios. 2) Human Alignment: Human preference annotations are provided for the five image-level dimensions, ensuring the benchmark's alignment with perceptual quality. 3) Valuable Insights: Beyond standard indoor settings, we analyze model performance variations across dimensions and investigate the disparity between indoor and real-world try-on scenarios. To foster the field of virtual try-on towards challenging real-world scenarios, VTBench will be open-sourced, including all datasets and evaluation algorithms.