From Design Specification to Manipulative UI: Auditing Specification-to-Behavior Failures in Agentic UI Generation
Abstract
Agentic UI-generation systems increasingly treat natural-language design specifications as executable inputs. This creates a new specification-to-behavior translation layer: design preferences that would traditionally be interpreted by designers and developers can be operationalized directly into interface behavior. We study whether this process can produce manipulative interaction affordances, including when downstream compliance requirements are provided. We evaluate 73 publicly available design specifications (70+ from VoltAgent, 3 from Google templates) across four choice-sensitive interface tasks: login, cookie consent, cookie preferences, and subscription, using two agentic UI-generation systems. We analyze both rendered interfaces and execution traces, with detected manipulative affordances subsequently validated by human reviewers. Across the evaluated generations, we observe recurring behavioral patterns including asymmetric visual salience, default-enabled options, bundled consent, and unequal interaction costs. We contribute to the following in this paper: (1) identification of textual references in design.md specifications that result in dark patterns; (2) trace-level analysis revealing systematic interpretation gaps including instances where ethical considerations are overridden ; and (3) classifying the failures into four related theses with one cross cutting property with empirical support explaining why specification-to-behavior gaps emerge in agentic execution. An exploratory intervention adding a generic GDPR-compliance instruction does not reliably eliminate these behaviors, suggesting that declarative compliance language alone may be insufficient to constrain specification-to-behavior translation. This work establishes specifications being auto-gamed in agentic UI generation as a distinct AI alignment challenge and argues that effective governance must move beyond static instruction writing toward trace-driven oversight and accountable checkpoints to avoid design instructions becoming ethically consequential.