What Makes Generated Content Functionally Legible? A VLM-Judged Ablation with Executable Ground Truth
Abstract
Image generation models make game art cheap to generate, but a game asset has a job: a player must see at a glance what each piece does. This property, which we call functional legibility, is exactly what generation pipelines quietly break, because the generator is told how each asset should look and nothing about what it means. Verifying legibility is just as hard, because the natural evaluator, a vision-language model, is an opinion rather than a measurement: easily fooled by its prompt, unstable across runs, and rarely held to any ground truth. We study both problems in a Match-3 game whose engine defines every asset's gameplay role in shipping code, so generation gets a specification and evaluation gets an executable answer key, and the two are the same object. On the evaluation side, a judge must prove it is looking before it may score, and no number is reported from a single scoring pass. Under this measurement, ablating the pipeline one mechanism at a time gives a clean attribution: legibility comes from grounding generation in gameplay roles, while the reference images and critic loop stacked on top change review effort, not legibility. Grounding is necessary but not sufficient: one theme defeats it entirely, and we report where and why instead of averaging the failure away.