How Memory Format and Structural Priors Shape an LLM Agent's Iterative Improvement on Chinese-Character Drawing
Abstract
Can an AI agent iteratively improve at a hard compositional task by accumulating memory, and does the format of that memory matter? We study this on programmatic Chinese-character drawing — 668 items judged by a blind human rater on a four-tier quality rubric — with five agent conditions whose matched contrasts separate memory format from an injected structural prior derived from stroke-skeleton data. We find that the structural prior is the dominant lever: injecting it raises the success rate (Perfect+Pass) by roughly 24 percentage points and unlocks the top quality tier, while memory format alone contributes little and yields no quality gains, even though each group’s memory store grew several times over during the experiment. Format does, however, shape downstream self-correction: it conditioned how effectively agents recovered their own failures on retry. We also observe that the memoryless control was the only condition to discover a degenerate shortcut — rendering glyphs with a system font — suggesting that accumulated procedural memory anchors an agent’s strategy search. For iterative, agent-driven applications, these results argue for grounding agents in verifiable domain structure first and treating memory design as a second-order choice whose main role is expressing, not achieving, quality