Budget Guard: Calibrated Resource Control for Long-Horizon LLM Agents
Jermaine He ⋅ Andrew Somerset ⋅ Aye Chan Myat Phyo ⋅ Laryn Qi ⋅ Charlotte Le ⋅ Ruizhe Li
Abstract
Long-horizon LLM agents can consume substantial amounts of tokens, time, and money before recognizing that a task is unlikely to succeed. Prior work on budget-aware agents shows the promise of early stopping, but agents' self-reported budget estimates are unreliable and cannot be used for decisions. We introduce Budget Guard, a lightweight, CPU-only wrapper for any LLM agent with two separate parts: an interval that estimates how many resources a task still needs, calibrated with conformal prediction so it is correct about as often as intended, and a feasibility gate that decides when to stop. We calibrate and evaluate the two separately. For the interval: across 19 agent–task combinations in the single-resource (token) environments (one split each), the agents' own estimates cover the truth only 2\% to 52\% of the time, while the guard's interval stays above 85\% coverage on 18 of 19; and is narrower than a feature-free baseline on two of three token environments (22–55\%), though wider on the third. For the gate, in a counterfactual replay of logged rollouts (not live deployment), we present a case study in Warehouse, a multi-resource environment (time, capacity, cost): the gate roughly halves the cost per successful run (\$1.61M to \$0.87M), but saves far less on single-resource token tasks (3–7\%), and adding the interval into the stop rule only makes the gate more conservative.
Chat is not available.
Successful Page Load