PainterBench: Evaluating Painterly Controllability in Generation and Verification
Abstract
For painters and illustrators, image generation models can produce compelling outputs, yet their ability to reliably follow precise painterly direction remains difficult to evaluate. We refer to this ability as painterly controllability. Although creative agency encompasses more than control alone, controllability is a necessary component: creators must be able to specify visual decisions and have them reliably realized. We introduce PainterBench, a dual-purpose and reference-based benchmark framework for painterly controllability in image generation and vision-language verification. PainterBench uses completed artworks as reproducible proxies for realized visual intent and measures fidelity along five dimensions derived from recurring concerns in painting practice: composition, value, palette, edge organization, and surface treatment. A verifier benchmark evaluates vision-language models using image variants with known dimension-specific orderings, while a generator benchmark compares image generation models under progressively richer text-only and multimodal specifications. Experiments reveal substantial variation across generators, control interfaces, painterly dimensions, and verifiers. PainterBench provides a reproducible framework for measuring painterly controllability as one technical foundation of creative agency in AI-assisted painting and illustration. The complete benchmark, including its datasets, evaluation protocols, and code, is publicly available at https://github.com/jeffreyliuster/painterbench.