No Consistent Now: Fractured Tool Snapshots and Unsafe LLM-Agent Commits in a Controlled Runtime
Abstract
Tool results can be individually truthful yet jointly describe no state that ever existed. We test this fractured-snapshot behavior in a deterministic simulator with globally comparable, episode-scoped snapshot tokens. The fixed, prompt-conditioned census contains 540 episodes: three checkpoints × 36 domain–template instances × five conditions, with one temperature-zero decode per episode. In each unguarded fractured condition, 107/108 episodes produced a protocol-unsafe commit effect; exposing snapshot tokens left this aggregate count unchanged under the evaluated prompt. A fail-closed guard then enforced the simulator's admissibility predicate, yielding 0/108 accepted unsafe effects by construction. Completion on one matched positive fixture was 107/108 both with and without the guard. Thus the behavioral result is a narrow observation about three checkpoints under one external prompt, while the guarded result is an implementation check rather than evidence of general agent reliability or production safety.