ArrivalBench: Agent-Generated Data Pipelines Are Correct Once and Wrong Under Time
Abstract
Benchmarks for agent-generated data work grade a pipeline by executing it once against a fixed snapshot. We show that such a check cannot discriminate temporal correctness even where it certifies nearly everything. Across eleven models from eight organisations it certifies 86–100% of produced pipelines, ten of the eleven at or above 97.9%; re-executing the same artifacts under adversarial but replayable delivery schedules separates them across 7.0% to 79.2% silent failure, an eleven- fold range the incumbent metric cannot resolve, though the shape of that range is two frontier models separating from a cluster rather than a graded ordering. Ar- rivalBench verifies what an agent left behind rather than what it did: the produced pipeline is re-executed after the agent stops, and its final state must equal a batch recomputation of the complete logical log under every schedule. The oracle is a recomputation, not a classifier, so a wrong answer and a crash are distinct verdicts, and the two are not equally costly: a crash surfaces at once and gets fixed, while a wrong revenue rollup is found a quarter later, after downstream consumers have built on it. First, the failures share a shape across organisational boundaries: in every unhinted arm the weakest ordering hazard scores above the weakest idem- potency hazard, so exactly-once is the weaker axis in every one, though in the two weakest models that margin means less because ordering itself has collapsed. Second, a prompt can relocate failure rather than remove it: a hazard warning cuts one model’s silent failure from 48.2% to 10.5% while raising its crash rate from 9.0% to 37.0%, improving all-in failure only from 51.0% to 44.0%. Of seven in- tervention arms five repair, one relocates, and one leaves the model no better than silence; only a grader separating a wrong answer from a crash tells which. Finally, a small human anchor points the other way: four practising engineers given the identical prompt on five of the tasks fail ordering where the models fail exactly- once. On seven to nine usable cells that is a dissociation, not a proof, and one model arm inverts on the same tasks, so the anchor bounds the two task-design readings rather than settling them. All eleven unhinted arms were independently re-run: rates move at most 5.9 points and the failing tasks largely recur.