When Language Models Talk Themselves Into Giving Up
Jayden Bai
Abstract
Self-improving agents adapt by several routes, including weight updates, few-shot prompting, and persistent memory they write for themselves. However, it is not yet clear whether a model can manage that memory and meaningfully improve towards completing a research task. We introduce a closed-loop harness in which a language model iteratively writes, executes, and revises its own Python solver against a synthetic symbolic-regression task, adapting only through text markdown files it authors and re-reads each iteration. Our four conditions vary where the model may write: an object-level search strategy, a meta-level set of operating instructions, both, or neither. Across two models (Claude Haiku 4.5 and Claude Sonnet 4.6) this produces roughly 1,200 solver programs. With $N=5$ repetitions per condition, the effects are directional rather than statistically significant, but the picture is consistent. Allowing the model to rewrite its own search strategy reduces the median test error by roughly 23--31%, while allowing it to rewrite its own reasoning instructions destabilizes the loop, doubling the rate of degenerate runs from $7$ of $20$ to $15$ of $20$. Most notably, failure is not solely attributed to poor search: in one run a model declared further improvement mathematically impossible at an error well above its own previously discovered best, then re-emitted an identical solver for $17$ consecutive iterations, locking in a result about $60\%$ worse in validation RMSE than one it had already found. We conclude that today's models can optimize a learning curve when the loop steers them, but cannot yet be trusted to steer their own research out of the box.
Chat is not available.
Successful Page Load