Looking Early at a Simulated Control Arm: Machine-Checked Anytime-Valid Bounds under Interim Looks, Continuation and Arm Selection
Abstract
We measure what three monitors of a running treatment effect do when the control arm is a patient simulator that can be queried at any time: a fixed-sample interval, a Lan-DeMets O'Brien-Fleming spending design, and two anytime-valid confidence sequences whose validity theorems are checked in Lean 4 against Mathlib. Four analyst behaviours that such a simulator invites are tested on a synthetic two-arm trial with six predictions sealed before the run (planned looks, every-patient looks, continuation past the planned maximum, best-of-k simulator reruns); two post-review cells add drift and a fitted simulator. The spending design's false-claim rate is compatible with nominal on its planned schedule and on a dense schedule it is re-planned for (its validity there rests on construction); it fails when its final boundary is reused past the planned maximum (18.1%) and when the best of 5 simulated arms is reported (19.0%). The betting sequence's false-claim rate is below 5% in every sealed cell, with every Wilson 95% interval under 5%; within the selection replications the first arm alone claims in 1.0% and the best of five in 3.3% (interval 2.6 to 4.2), a post-hoc run at 20 reruns reaches 12.8%, and a δ/k union bound (δ/(2k) per side) gives 0.2%. A fixed-sample interval at α/5 still claims falsely 10.2% under every-patient looks. At the planned size the betting and mixture half-widths are 1.41 and 8.8 times the fixed-sample half-width. A biased, drifting or poorly fitted simulator misleads the fixed-sample and betting intervals, the sharper paired interval most of all, while the mixture keeps coverage only by being too wide to reject anything; none of the intervals tested repairs a wrong world model. Graded after the run under a stated rule, four sealed predictions held as written and two failed in a clause.