Latency Limits on Safety Guarantees for Language Model Controllers
Abstract
Language models are being put in charge of systems that act. When the action is a command to a physical plant, the model’s inference time is not just overhead, it is time the plant keeps moving. We ask whether the worst-case safety guarantees of Hamilton–Jacobi reachability survive that, and report that in our setting they do not. Treating inference latency as a zero-order-hold delay, we derive the disturbance bound it induces. The relevant quantity is the diameter of the reachable control effect set, not its radius. We then compute the backward reachable tube across delay. The safe set erodes slowly to 0.04 s, turns sharply by 0.10 s, and is empty from 0.130 s onward, twelve times below the 1.5543 s CPU latency we measure, and still below the 0.1393 s we measure on a consumer GPU, where even the shortest response the pipeline can produce is already past the threshold. The Hamiltonian’s own dimensionless scale overstates this budget threefold. A second component points the other way: capping a latent safety margin’s Lipschitz constant threefold preserved recall and cut false alarms 5.4×, a strict improvement rather than a trade off, though the monitor’s discrimination is directional at this N once sequence-level clustering is accounted for. These are characterized boundaries, not tuning gaps.