What Self-Evolution Leaves Behind: Observing Skill Formation Beyond Endpoint Scores
Abstract
Self-evolving agents do not merely change task performance. Their interactions leave behind memories, procedures, and code that can organize capability in qualitatively different ways, yet endpoint scores collapse these structures into a scalar outcome. We ask what capability structure, if any, self-evolution forms. We treat skill formation as the emergence of a reusable structure supported by invocation awareness, executable grounding, measurable effect, and observable boundaries. To make these properties visible, we introduce a rule-grounded adversarial assay built from an executable referee, fixed diagnostic opponents, replayable trajectories, constructed pressure states, and artifact interventions. We instantiate the assay in Gomoku, a model organism in which actions, outcomes, trajectories, and generated artifacts can be checked and manipulated. Across 268 games with Reflexion-style textual accumulation, Voyager-style executable skill libraries, and DGM-style code evolution, we observe distinct forms of skill formation that endpoint outcomes conceal. Reflexion memories increased pressure-state passes by two or three out of 22 but produced 0/60 natural-game wins and 0/60 successful must-block responses. Voyager ablations changed threat-aware wins from 7/10 to 3/10 and live-four creation from 5/5 to 0/5 even though every variant scored 0/5 against the strongest opponent. The retained DGM path improved from 15/22 to 19/22 pressure-state passes through archived decision-procedure revisions. Together, these cases expose distinct transformations from retaining strategic content, to grounding it in executable procedures, to coordinating procedures into broader decision structures, while locating the further boundary at the autonomous formation of predictive search.