Documentation Helps, But Gaps Remain: Diagnosing LLM Generation for Evolving Diagram DSLs
Abstract
Large language models (LLMs) are increasingly used to generate executable artifacts in evolving diagram domain-specific languages (DSLs). In these domains, existing documentation improves render rates but fails to eliminate the render-constraint gap: code that renders successfully while violating task-critical structural constraints. We introduce DIAGRAM-FAIL-BENCH, a 630-task Mermaid benchmark across 21 diagram types with stale syntax targets, required facets, deterministic checks, and semantic checkpoints. We also use RA-RAG as a controlled, verifier-guided repair intervention combining facet-aware retrieval, syntax contracts, deterministic validation, structured diagnostics, and repair-dedicated retrieval. Experiments across three model families show that raw execution-feedback repair eliminates most of the gap, while RA-RAG further improves executable machine success. Furthermore, we audit verifier implementations and constraint schemas, report human-checked semantic audits across methods, and stress-test modern coding agents on a frozen 70-task subset: black-box verifier iterations raise the final executable success rate of the same model from 51.43% to 60.00%, while the task-aligned RA-RAG reference reaches 95.71%. We report the latter as descriptive stress-testing evidence rather than a causal mechanism estimate, as it changes both model and framework.