Evaluating Interactivity Judges for Generated Web User Interfaces
Abstract
Large language models are increasingly used to generate web user interfaces in response to user requests. Such WebUIs are created either by directly prompting LLMs or through coding assistants. However, while frontier LLMs are increasingly capable of generating WebUIs of high aesthetic quality, functional interactive components remain challenging. During the evaluation or post-training of LLMs for WebUI generation, the quality of interactive elements is quantified using an interactivity judge. Interactivity judges are themselves LLM-based agents that typically operate on generated WebUIs through browser interactions and screenshot analysis, producing a numeric quality score for the interactive elements on the page. We introduce InteractivityJudgeBench, the first benchmark for evaluating WebUI interactivity judges consisting of 413 human-verified web artifacts across diverse application categories. For each prompt, we provide a working implementation and up to five variants with increasingly broken interactive elements. Each variant is accompanied by a structured manifest documenting exactly which elements were broken and how, providing objective and human-verifiable ground truth. We propose four relative ordering metrics that measure how well a judge distinguishes adjacent quality levels at coarse, fine, gap-weighted, and full-ranking granularity. Next, we benchmark a range of non-agentic and agentic interactivity judges. We further propose an agentic judge that addresses the limitations of prior methods and outperforms existing agentic baselines across all proposed metrics. However, the more surprising finding is that a simple rubric-based non-agentic judge matches or exceeds agentic judges while being significantly faster, suggesting that agentic setups may be unnecessary for WebUI interactivity evaluation.